Detection method and device based on visual model

Through the dynamic group channel pruning method based on structural dependence, unimportant channels in the underground coal mine vision model are evaluated and pruned, which solves the problem of lightweight and high-precision detection of the vision model in a resource-constrained environment, and realizes high-reliability real-time detection of the environment and personnel status in the underground coal mine.

CN120635764APending Publication Date: 2025-09-12TIANDI TECH CO LTD BEIJING TECH RES BRANCH +1

Patent Information

Application Number
CN202510577803.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-09-12

Smart Images

  • Figure CN120635764A_ABST
    Figure CN120635764A_ABST
Patent Text Reader

Abstract

The invention provides a detection method and device based on a visual model, and the method comprises the steps: obtaining an image sample training set of any visual model in a coal mine video monitoring system; according to a dependency relationship between elements of each layer in the visual model, each element is divided into a plurality of groups, and the elements are layer input or layer output of each layer; performing iterative training on the visual model by adopting the image sample training set, and in each round of iterative training, determining an importance index value of the intra-group channel of at least one group according to the weight of the element in the at least one group, the weight gradient and the attention weight of the corresponding intra-group channel, according to the importance index value of the intra-group channel of the at least one group, updating the model parameters of the visual model; the trained visual model is used for detecting the to-be-detected target image in the coal mine video monitoring system, the detection result of the target image is obtained, and the detection accuracy can be kept while the structure of the visual model is compressed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of coal mines, and in particular to a detection method and device based on a visual model. Background Art

[0002] As an important part of the safety monitoring system, the coal mine video surveillance system is of great significance for real-time monitoring of the underground environment and personnel status.

[0003] Related technologies use various visual models in coal mine visual monitoring systems to monitor the underground environment and personnel status in real time. However, coal mines have limited space, and there are strict restrictions on the computing power, storage space, and energy consumption of visual models. Furthermore, real-time processing and high reliability are required. Therefore, within limited computing resources, visual models need to be lightweight and provide high-precision detection. Summary of the Invention

[0004] The present invention aims to solve one of the technical problems in the related art at least to a certain extent.

[0005] To this end, the first object of the present invention is to propose a detection method based on a visual model to achieve compressing the visual model structure while maintaining the accuracy of detection.

[0006] The second object of the present invention is to provide a detection device based on a visual model.

[0007] A third object of the present invention is to provide an electronic device.

[0008] A fourth object of the present invention is to provide a computer-readable storage medium.

[0009] A fifth object of the present invention is to provide a computer program product.

[0010] To achieve the above objectives, a first embodiment of the present invention proposes a detection method based on a visual model, comprising:

[0011] Obtain an image sample training set of any visual model in a coal mine video monitoring system, wherein the image sample training set includes a detection object and detection information of the detection object, and the visual model is a network model;

[0012] Dividing the elements of each layer in the visual model into a plurality of groups according to dependencies between the elements, wherein the elements are layer inputs or layer outputs of each layer;

[0013] Iteratively training the visual model using the image sample training set, and in each round of iterative training, determining an importance index value of a channel within at least one of the groups based on a weight, a weight gradient, and an attention weight of an element in the group, and updating model parameters of the visual model based on the importance index value of the channel within the group, wherein the importance index value is used to indicate the importance of the channel within the group in the corresponding group;

[0014] The trained visual model is used to detect the target image to be detected in the coal mine video monitoring system to obtain the detection result of the target image.

[0015] To achieve the above objectives, a second embodiment of the present invention provides a detection device based on a visual model, comprising:

[0016] An acquisition module is used to acquire an image sample training set of any visual model in a coal mine video monitoring system, wherein the image sample training set includes a detection object and detection information of the detection object, and the visual model is a network model;

[0017] a grouping module, configured to divide the elements of each layer in the visual model into a plurality of groups according to dependencies between the elements of each layer, wherein the elements are layer inputs or layer outputs of each layer;

[0018] a training module, configured to iteratively train the visual model using the image sample training set, and in each round of iterative training, determine an importance index value of a channel within at least one of the groups based on the weights, weight gradients, and attention weights of the elements in at least one of the groups and the corresponding channel within the group, and update model parameters of the visual model based on the importance index value of the channel within the group, wherein the importance index value is used to indicate the importance of the channel within the group in the corresponding group;

[0019] The detection module is used to detect the target image to be detected in the coal mine video monitoring system using a trained visual model to obtain a detection result of the target image.

[0020] To achieve the above-mentioned object, a third embodiment of the present invention provides an electronic device, comprising:

[0021] at least one processor; and

[0022] a memory communicatively connected to the at least one processor; wherein,

[0023] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method according to the first aspect.

[0024] In order to achieve the above-mentioned purpose, an embodiment of a fourth aspect of the present invention proposes a computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable the computer to execute the method described in the first aspect.

[0025] In order to achieve the above-mentioned purpose, a fifth embodiment of the present invention proposes a computer program product, including a computer program, which implements the method described in the first aspect when executed by a processor.

[0026] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0028] Figure 1 A flowchart of a detection method based on a visual model provided by an embodiment of the present invention;

[0029] Figure 2 A schematic diagram of the structure of a visual model provided by an embodiment of the present invention;

[0030] Figure 3 A schematic flow chart of another detection method based on a visual model provided by an embodiment of the present invention;

[0031] Figure 4 A schematic flow chart of another detection method based on a visual model provided by an embodiment of the present invention;

[0032] Figure 5 A schematic structural diagram of a detection device based on a visual model provided by an embodiment of the present invention;

[0033] Figure 6 A schematic structural diagram of an electronic device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0034] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.

[0035] As an important part of the safety monitoring system, the coal mine video surveillance system is of great significance for real-time monitoring of the underground environment and personnel status.

[0036] Related technologies use various visual models in coal mine visual monitoring systems to monitor the underground environment and personnel status in real time. However, coal mines have limited space, and there are strict restrictions on the computing power, storage space, and energy consumption of visual models. Furthermore, real-time processing and high reliability are required. Therefore, within limited computing resources, visual models need to be lightweight and provide high-precision detection.

[0037] Model lightweighting refers to the use of technical means to reduce the parameter size and computational complexity of machine learning models so that they can run efficiently on resource-constrained devices. Related technologies include model lightweighting methods such as model pruning, parameter quantization, knowledge distillation, and architecture optimization. Model pruning reduces model size by removing redundant connections and neurons; parameter quantization reduces storage requirements by reducing the precision of weights; knowledge distillation achieves performance compression and transfer by training small models to mimic the behavior of large models; and architecture optimization focuses on designing more efficient network structures, such as mobile networks and efficient convolutional modules.

[0038] Model pruning is primarily achieved through two methods: unstructured pruning and structured pruning. While unstructured pruning can effectively reduce model parameters, it typically requires numerous fine-tuning training steps to recover performance losses caused by pruning, which consumes training resources and prolongs training time. Structured pruning, on the other hand, preserves the overall structure of the model by pruning entire structurally significant parts, such as neurons, channels, or layers, thereby achieving model compression while maintaining good performance. However, structured pruning has limitations in its application across different types of models and lacks sufficient versatility, making it difficult to apply broadly across multiple model architectures.

[0039] Furthermore, structured pruning methods primarily rely on a single metric or a combination of a few composite metrics when evaluating channel importance, failing to comprehensively and accurately measure the actual contribution of each channel. These methods typically judge channel importance based on weight size or activation value, but a single weight or activation metric cannot fully reflect the performance of a channel under different input data or training stages. Furthermore, most evaluation methods are static and fail to dynamically adjust channel importance to adapt to changes in the training process or input data. They ignore gradient information and task-specific requirements, and are unable to fully capture channel importance, thus limiting the optimization of pruning decisions and improving overall model performance.

[0040] In response to the above problems, the present invention proposes a detection method based on a visual model, which compresses the visual model structure while maintaining the accuracy of detection.

[0041] In this paper, a dynamic group channel pruning method based on structural dependencies is used to compress the visual model structure. This pruning method achieves structured pruning of the model structure with a specified proportion and structure by building channel structure dependencies, achieving versatility while maintaining model accuracy as much as possible.

[0042] In this paper, a channel importance assessment method based on the correlation between layer input and output structures is used to identify the channels that contribute least to model performance. This method utilizes the correlation between layer input and output structures and uses multiple importance assessment metrics to evaluate channel importance. This ensures that pruned channels have a minimal impact on the model output, thereby reducing model parameters and computational complexity while maintaining the accuracy of the model output results.

[0043] The following describes a detection method and apparatus based on a visual model according to an embodiment of the present invention with reference to the accompanying drawings.

[0044] Figure 1 A flowchart of a detection method based on a visual model provided by an embodiment of the present invention.

[0045] like Figure 1 As shown, the detection method based on the visual model includes the following steps:

[0046] Step 101: Obtain an image sample training set of any visual model in a coal mine video monitoring system.

[0047] The image sample training set includes the detection object and the detection information of the detection object. For example, an image of an underground miner includes the miner (detection object) and the miner's location information, and / or the miner's helmet (detection object) and the location information of the miner's helmet.

[0048] Among them, the visual model is a network model.

[0049] In some embodiments, the visual model can be any visual model in the coal mine video monitoring system, such as an underground vehicle detection model, an underground personnel safety helmet detection model, etc.

[0050] Step 102: Divide the elements into multiple groups according to the dependency relationship between the elements in each layer of the visual model.

[0051] Among them, the elements are the layer input or layer output of each layer.

[0052] In some embodiments, the layer inputs and layer outputs of all layers in the visual model can be divided into multiple groups (g1, g2, ..., g n), where each group includes mutually dependent layer inputs and / or layer outputs (e.g., is the first layer input, is the first layer output, is the second layer input).

[0053] It should be noted that in the present invention, the basis for determining whether a dependency relationship exists between layer inputs and layer outputs of different layers in a visual model is whether the layer input and layer output channels of the same layer have a one-to-one correspondence, or whether the layer input and layer output channels of different layers are the same channel. If the layer input and layer output channels of the same layer have a one-to-one correspondence, or the layer input and layer output channels of different layers are the same channel, then a dependency relationship exists between the two; otherwise, no dependency relationship exists.

[0054] As an example, in a neural network, the ReLU (Rectified Linear Unit) layer is a nonlinear activation function layer used to increase the nonlinear ability of the network. Since the ReLU layer performs nonlinear transformation on the input data element by element, the channels of the layer input and layer output of the ReLU layer are one-to-one corresponding; Batch Normalization (BN) The Batch Normalization (Batch Normalization) layer is used to normalize each channel of the input data to accelerate the training process and improve the stability of the model. Since the BN layer operates channel by channel, the channels of its layer input and layer output are also one-to-one corresponding. The Add layer is used to add two or more input tensors element-wise. To ensure that the addition operation can be performed correctly, all input tensors of the Add layer must have the same shape, including the number of channels. Therefore, the channels of the layer input and layer output of the Add layer are also one-to-one corresponding. The convolution layer generates an output feature map by convolving the input data with the convolution kernel. The number of input channels of the convolution layer is determined by the number of channels of the input data, while the number of output channels is determined by the number of convolution kernels. Therefore, the channels of the layer input and layer output of the convolution layer are usually not one-to-one corresponding. Each neuron in the fully connected layer (also called the dense layer) is connected to all neurons in the previous layer to map the features to the final output space. The input and output of the fully connected layer are flattened vectors. The channels of the layer input and layer output of the fully connected layer are usually not one-to-one corresponding, but are adjusted according to the needs of the network design. Therefore, if the i-th layer in the structure of the visual model is a ReLU layer, or a BN layer, or an Add layer, the channels of the layer input and layer output of the i-th layer are one-to-one corresponding. At this time, the layer input of the i-th layer and the layer output of the i-th layer have a dependency relationship, and the two are in the same group.

[0055] As another example, if the $i$-th layer and the $j$-th layer in the structure of the vision model are connected, where $i < j$, then the channels of the layer output of the $i$-th layer and the channels of the layer input of the $j$-th layer are the same channels. At this time, there is a dependency relationship between the layer output of the $i$-th layer and the layer input of the $j$-th layer, and the two are in the same group.

[0056] Step 103: Use the image sample training set to perform iterative training on the vision model. In each round of iterative training, determine the importance index values of the intra-group channels of at least one group according to the weights, weight gradients of the elements in at least one group, and the attention weight values for the corresponding intra-group channels, and update the model parameters of the vision model according to the importance index values of the intra-group channels of at least one group.

[0057] The importance index value is used to indicate the importance degree of the intra-group channel in the corresponding group.

[0058] In some embodiments, in each round of iterative training, use the image sample training set to perform iterative training on the vision model corresponding to this round of iterative training, and obtain the weights and weight gradients of each layer element in the vision model in this round of iterative training; where the vision model corresponding to the first round of iterative training is the vision model; for any target group including multiple elements, determine the importance index value of each intra-group channel of the target group in this round of iterative training according to the weights, weight gradients of each element in the target group in this round of iterative training, and the attention weight values of each element in the target group for the corresponding intra-group channels in this round of iterative training; update the model parameters of the vision model corresponding to this round of iterative training according to the importance index values of each intra-group channel of each target group in this round of iterative training to obtain the vision model corresponding to the next round of iterative training.

[0059] Among them, the target group includes multiple elements, each element has the same number of channels, and the channels of each element are either in one-to-one correspondence or the same channel.

[0060] The intra-group channels of the target group refer to the non-repeated channels of each element in the target group.

[0061] In some embodiments, for any target group including multiple elements, determine the importance index value of an element for a target channel in this round of iterative training according to the weight, weight gradient of the element in this round of iterative training, and the attention weight value of the element for the corresponding intra-group channel in this round of iterative training; where the target channel is any intra-group channel corresponding to the element; accumulate the importance index values of each element in the target group for the target channel in this round of iterative training to obtain the importance index value of the target channel in this round of iterative training.

[0062] As an example, assuming that each element in the target group has K channels, the importance index value of any element in the target group to the channel k in the group in any round of iterative training is Imp(k,f), where f represents the element and k represents the channel k. Then the importance index value of the channel k in the group in this round of iterative training is obtained by accumulating the importance index value of at least one element to the channel k in the group in this round of iterative training, that is, Imp(k) = ∑ f Imp(k,f). The importance index value Imp(k,f) of any element in the target group to channel k in any round of iterative training is determined based on the weights and weight gradients of each element in the target group in this round of iterative training, as well as the attention weight of each element on channel k in the group in this round of iterative training.

[0063] In some embodiments, for any target group including multiple elements, the weight of the target element in each channel within the target group in the current round of iterative training is determined according to the weight of the target element in the current round of iterative training; wherein the target element is any element in the target group; the ratio between the weight of the target element in the current round of iterative training and the sum of the weights of each channel within the target group in the current round of iterative training is determined as the weight of the target element on the target channel in the current round of iterative training; the weight gradient of the target element in the current round of iterative training for each channel within the target group is determined according to the weight gradient of the target element in the current round of iterative training; the weight gradient of the target element in the current round of iterative training for the target channel is determined according to the weight gradient of the target element in the current round of iterative training for the target channel. The ratio of the degree to the sum of the weight gradients of the target element on each channel within the target group in this round of iterative training is determined as the weight gradient of the target element on the target channel in this round of iterative training; the ratio of the attention weight of the target element on the target channel in this round of iterative training to the sum of the attention weights of the target element on each channel within the target group in this round of iterative training is determined as the weighted attention weight of the target element on the target channel in this round of iterative training; based on the weight of the target element on the target channel in this round of iterative training, the weight gradient of the target element on the target channel in this round of iterative training and the weighted attention weight of the target element on the target channel in this round of iterative training, the importance index value of the target element to the target channel in this round of iterative training is determined.

[0064] It should be noted that each element (i.e., each layer input / layer output) includes at least one layer input vector / layer output vector, and each element (i.e., each layer input / layer output) may have one or more layer input weights / layer output weights.

[0065] As an example, Figure 2 A schematic diagram of the structure of a visual model provided by an embodiment of the present invention. Figure 2Each circle in represents a channel, each channel corresponds to a layer input vector / layer output vector, and each line represents a layer input weight / layer output weight. Figure 2 As shown, the visual model has 3 layers, and all layers in the model input All layer outputs in, is the first layer input, is the first layer output, is the second layer input, is the second layer output, is the third layer input, The third layer output.

[0066] Figure 2 In the figure, the first column of circles represents the first layer input The second column of circles represents the output of the first layer Channel and second layer input Channel (the second layer is connected to the first layer, the second layer input The channel and the first layer output The channels are the same channels), the third column of circles represents the second layer output Channels and third layer input Channel (the third layer is connected to the second layer, the second layer output The channel and the third layer input The channels are the same channels), the fourth circle represents the output of the third layer channel.

[0067] pass Figure 2 It can be seen that the first layer input There are 3 channels, each channel corresponds to a layer input vector, the first layer input There are 9 lines, indicating the first layer input There are 9 layers of input weights; the first layer outputs There are 3 channels, each channel corresponds to a layer output vector, the first layer output There are 9 lines (the lines on the right side of the second circle represent the layer output weights of the first layer, and the lines on the left side represent the layer input weights of the second layer), indicating that the output of the first layer There are 9 layers of output weights; the second layer input There are 3 channels, each channel corresponds to a layer input vector, the second layer input There are 6 lines, indicating the second layer input There are 6 layers of input weights; the second layer outputs There are also 2 channels, each channel corresponds to a layer output vector, the second layer output There are 6 lines (the lines on the right side of the third circle represent the output weights of the second layer, and the lines on the left side represent the input weights of the third layer), indicating that the output of the second layer There are 6 layers of output weights; the third layer input There are 2 channels, each channel corresponds to a layer input vector, the third layer input There are 6 lines, indicating the third layer input There are 6 layers of input weights; the third layer outputs There are also 3 channels, each channel corresponds to a layer output vector, the third layer output There are 6 lines, indicating the output of the third layer There are 6 layers with output weights.

[0068] For the first layer input Assume that the first column of circles (the first layer input The first circle from top to bottom is channel 1, the second circle is channel 2, and the third circle is channel 3. The first layer input Among the 9 layer input weights, the layer input weights represented by the 3 lines connected to the first circle are the weights of the first layer input on channel 1, the layer input weights represented by the 3 lines connected to the second circle are the weights of the first layer input on channel 2, and the layer input weights represented by the 3 lines connected to the third circle are the weights of the first layer input on channel 3. The first layer output Second layer input Second layer output The third layer input The third layer output Similarly, I will not go into details here.

[0069] Each weight has a corresponding weight gradient, for the first layer input Assume that the first column of circles (the first layer input The first circle from top to bottom is channel 1, the second circle is channel 2, and the third circle is channel 3. The first layer input Among the weight gradients corresponding to the 9 layer input weights, the weight gradient corresponding to the layer input weight represented by the 3 lines connected to the first circle is the weight gradient of the first layer input with respect to channel 1, the weight gradient corresponding to the layer input weight represented by the 3 lines connected to the second circle is the weight gradient of the first layer input with respect to channel 2, and the weight gradient corresponding to the layer input weight represented by the 3 lines connected to the third circle is the weight gradient of the first layer input with respect to channel 3. The first layer output Second layer input Second layer output The third layer input The third layer output Similarly, I will not go into details here.

[0070] For each element in the above target group, there are K channels. The importance index value of any element in the target group to the channel k in the group in any round of iterative training is Imp(k,f), where f represents the element and k represents the example of channel k. Optionally, the importance index value of any element in the target group to the channel k in any round of iterative training is Imp(k,f)=α*Weight fk +β*Grad fk +γ*Att fk ,

[0071] Among them, α, β, and γ represent coefficients; Weight fk Weight represents the weight of element f on channel k in the group in this round of iterative training. fk The weight of each channel in the target group is determined based on the element f in this round of iterative training; Grad fk Represents the weight gradient of element f on channel k in the group during this round of iterative training, Grad fk Determine the weight gradient of each channel within the target group based on the element f in this round of iterative training; Att fk Indicates the weighted attention value of element f on channel k in the group in this round of iterative training, Att fk The attention weight of each channel within the target group is determined based on the element f in this round of iterative training.

[0072] Optionally, Weight fk 、Grad fk 、Att fk It is calculated using the following formula:

[0073]

[0074] Among them, when the element f is the layer input and the corresponding layer is a convolutional layer or a fully connected layer, Represents the weight of the channel in any group of the target group in the layer input f in this round of iterative training; when the element f is the layer output and the corresponding layer is a convolutional layer or a fully connected layer, The representation layer output f represents the weight of the channel within any group of the target group in this round of iterative training.

[0075] From the above formula, we can see that when the element f is the layer input and the corresponding layer is a convolutional layer or a fully connected layer, Weight fk It is the sum of the absolute values ​​of the weights of the channel k in the group of the layer input f in this round of iterative training; when the element f is the layer output and the corresponding layer is a convolutional layer or a fully connected layer, Weight fkIt is the sum of the absolute values ​​of the weights of the layer output f on the channel k in the group in this round of iterative training; when the element f is other, Weight fk is 0. And in the preliminary calculation, the Weight fk After that, it is necessary to weight the weight of each channel within the target group based on the element f in this round of iterative training, and weight the element f in this round of iterative training on the channel k within the group fk The sum of the weights of each channel within the target group in this round of iterative training with element f The ratio between them is determined as the weight of element f on the target channel in this round of iterative training. fk .

[0076]

[0077] Among them, when the element f is the layer input and the corresponding layer is a convolutional layer or a fully connected layer, Represents the weight of the channel in any group of the target group in the layer input f in this round of iterative training; when the element f is the layer output and the corresponding layer is a convolutional layer or a fully connected layer, The representation layer output f represents the weight of the channel within any group of the target group in this round of iterative training.

[0078] From the above formula, we can see that when the element f is the layer input and the corresponding layer is a convolutional layer or a fully connected layer, Grad fk It is the sum of the absolute values ​​of the weight gradients of the layer input f with respect to the channel k in the group in this round of iterative training; when the element f is the layer output and the corresponding layer is a convolutional layer or a fully connected layer, Grad fk It is the sum of the absolute values ​​of the weight gradients of the layer output f with respect to the channel k in the group in this round of iterative training; when the element f is other, Grad fk is 0. And in the preliminary calculation, Grad fk After that, it is necessary to weight the element f in this round of iterative training on the weight gradient of each channel within the target group, and the weight Grad of the element f in this round of iterative training on the channel k within the group fk The sum of the weights of each channel within the target group in this round of iterative training with element f The ratio between them is determined as the weight Grad of element f on the target channel in this round of iterative training. fk .

[0079]

[0080] From the above formula, we can see that when element f is the layer output and the corresponding layer is the convolutional layer, Att fkis the attention weight of the corresponding layer of f in channel k; when element f is other, Att fk The attention weight is determined by the sigmoid layer output in the channel attention module added under the convolutional layer of the visual model. The channel attention module includes a global average pooling layer, two fully connected layers, and a sigmoid layer.

[0081] And in the preliminary calculation, we get Att fk After that, it is necessary to weight the attention weight of each channel in the target group based on the element f in this round of iterative training, and the attention weight Att of the element f on the channel k in the group in this round of iterative training is fk The sum of the attention weights of each channel within the target group in this round of iterative training with element f The ratio between them is determined as the weight Att of element f on the target channel in this round of iterative training fk .

[0082] In some embodiments, for any target group, candidate channels are determined from the intra-group channels of the target group based on the correspondence between each element in the target group and the intra-group channels of the target group and the weight mask value of each element in the target group in this round of iterative training; wherein, the weight mask values ​​of the elements corresponding to the candidate channels in this round of iterative training are not the first set values, and the weight mask values ​​of each element in the visual model in the first round of iterative training are all the second set values; according to the importance index values ​​of the candidate channels, the candidate channels are sorted to select a target number of target channels from the sorted candidate channels; wherein, the target number is determined based on the iterative update ratio corresponding to this round of iterative training and the total number of intra-group channels of the target group; the weights and weight mask values ​​of the elements corresponding to the selected target number of target channels are set to the first set values ​​to obtain the visual model corresponding to the next round of iterative training.

[0083] The first set value and the second set value are two different values. For example, the first set value may be 0, and the second set value may be 1.

[0084] The target number may be the product of the iterative pruning ratio corresponding to the current round of iterative training and the total number of channels in the target group.

[0085] As an example, in each round of iterative training, for any target group g i , ignore the channels with a Mask of 0 in the group (i.e., the input weight or output weight Mask of the layer corresponding to the channel is 0), sort the remaining channels from small to large according to the importance index value Imp(k), and select the top Ratio Pruned *c_g iChannels, and set all layer input weights and Masks and layer output weights and Masks in the group corresponding to these channels to 0. Pruned is the iterative pruning ratio corresponding to the corresponding round of iterative training, c_g i Group the target g i The total number of channels in the group.

[0086] In some embodiments, in any round of iterative training, if the weight mask value of any element in the previous round of iterative training is the first set value, the weight gradient of the element in this round of iterative training is set to the first set value, and no update operation is performed on the weight of the element in this round of iterative training.

[0087] In some embodiments, whether to stop iterative training can be determined based on the number of iterative training times.

[0088] As a possible implementation method, after obtaining the visual model corresponding to the next round of iterative training, the number of training times corresponding to the previous round of iterative training can be increased by a third set value to obtain the number of training times corresponding to the current round of iterative training; wherein the number of training times corresponding to the first round of iterative training is the fourth set value; and when the number of training times corresponding to the current round of iterative training is not less than the set number of training times, the iterative training is stopped.

[0089] As an example, assuming that the number of training times is set to 5, the number of training times corresponding to the first round of iterative training is 1 (the fifth set value), and after each round of iterative training is completed, the number of training times corresponding to the previous round of iterative training is increased by 1 (the fourth set value) to obtain the number of training times corresponding to the current round of iterative training. The number of training times corresponding to the second round of iterative training is 1+1=2, 2 is lower than 5, and the iterative training continues. The number of training times corresponding to the third round of iterative training is 2+1=3, 3 is lower than 5, and the iterative training continues. The number of training times corresponding to the fourth round of iterative training is 3+1=4, 4 is lower than 5, and the iterative training continues. The number of training times corresponding to the fifth round of iterative training is 4+1=5, 5 is not lower than 5, and the iterative training is stopped.

[0090] In some embodiments, a trained visual model may be obtained based on the corresponding visual model of the last round of iterative training.

[0091] In some embodiments, the attention weight is determined based on the output value or preset value of the channel attention module added under the convolutional layer of the visual model, so that a trained visual model can be obtained by removing the added channel attention module and the weight mask value of the first set value in the visual model corresponding to the last round of iterative training.

[0092] Step 104 : Use the trained visual model to detect the target image to be detected in the coal mine video monitoring system to obtain a detection result of the target image.

[0093] The visual model-based detection method of the embodiment of the present invention obtains an image sample training set of any visual model in a coal mine video surveillance system; divides each element into multiple groups based on the dependency relationship between each layer element in the visual model, wherein the element is the layer input or layer output of each layer; iteratively trains the visual model using the image sample training set, and in each round of iterative training, for any target group including multiple elements, updates the model parameters of the visual model based on the weight, weight gradient, and attention weight of each element in the target group to the corresponding intra-group channel in the current round of iterative training; uses the trained visual model to detect the target image to be detected in the coal mine video surveillance system, and obtains the detection result of the target image. Thus, by dividing the groups based on the dependency relationship between the layer input and layer output of each layer in the visual model, and determining the importance index value of the intra-group channel of at least one group, unimportant channels are pruned based on the importance index value of the intra-group channel of at least one group, thereby reducing model parameters and computational complexity, and achieving universality while maintaining model accuracy as much as possible.

[0094] This embodiment provides another detection method based on a visual model. Figure 3 A schematic flow chart of another detection method based on a visual model provided by an embodiment of the present invention.

[0095] like Figure 3 As shown, the detection method based on the visual model may include the following steps:

[0096] Step 301: Obtain an image sample training set of any visual model in a coal mine video monitoring system.

[0097] Step 302: Based on the structure of the visual model, construct the dependency matrix DM of the visual model according to rules 1, 2, and 3. ij .

[0098] in,

[0099] DM ij =0 means no dependency, DM ij =1 indicates a dependency relationship, L indicates the number of layers of the visual model, 0<i≤2L, 0<j≤2L,

[0100] Dependency Matrix DM ij Used to indicate the dependencies between elements of each layer in the visual model.

[0101] Rule 1 is that if the $i$-th layer and the $j$-th layer in the structure of the vision model are connected, where $i < j$, then $DM$ (i+L)j $= 1$, $DM$ j(i+L) $= 1$,

[0102] Rule 2 is that if the input and output channels of the $i$-th layer in the structure of the vision model are in one-to-one correspondence, then $DM$ (i+L)i $= 1$, $DM$ i(i+L) $= 1$,

[0103] Rule 3 is that the connection is transitive. If $DM$ (i+L)j $= 1$ and $DM$ j(k+L) $= 1$, then $DM$ (i+L)(k+L) $= 1$.

[0104] In some embodiments, a layer is a basic building block of a deep learning model. Each layer generates an output from an input through specific mathematical operations, and multiple layers are combined to form a complex model. The input and output of each layer are respectively defined as Assume the model has $L$ layers. All layer inputs in the model All layer outputs

[0105] Initialize an input-output layer dependency matrix $DM = 0$<00,00049>, where the meaning of $DM$ ij in the matrix is as follows:

[0106]

[0107] $DM$ ij $= 0$ indicates no dependency, and $DM$ ij $= 1$ indicates a dependency;

[0108] When constructing the dependency matrix $DM$ ij of the vision model based on the structure of the vision model, it is constructed according to the following rules:

[0109] Rule 1: Assume that the $i$-th layer and the $j$-th layer in the model structure are connected, where $i < j$, then $DM$ (i+L)j $= 1$, $DM$ j(i+L) $= 1$;

[0110] Rule 2: Assume that the input and output channels of the $i$-th layer in the model structure are in one-to-one correspondence (such as ReLU layer, BN layer, Add layer), then $DM$ (i+L)i $= 1$, $DM$ i(i+L) $= 1$;

[0111] Rule 3: The connection is transitive. If $DM$ (i+L)j $= 1$ and $DM$ j(k+L) $= 1$, then $DM$ (i+L)(k+L) $= 1$.

[0112] As an example, for Figure 2 The visual model shown, all layers in the model input All layer outputs Assume that the first layer is a ReLU layer, or a BN layer, or an Add layer, and the second and third layers are fully connected layers or convolutional layers. Based on the structure of the visual model, the dependency matrix DM of the visual model is constructed. ij as follows:

[0113]

[0114] Among them, DM 14 =1, DM 41 =1 means that the first layer input and the first layer output have a dependency relationship (according to Rule 2, the first layer is a ReLU layer, or a BN layer, or an Add layer), DM 24 =1, DM 42 =1 means that the second layer input has a dependency relationship with the first layer output (according to rule 1, the second layer is connected to the first layer), DM 35 =1, DM 53 =1 means that the third layer input has a dependency relationship with the second layer output (according to rule 1, the second layer is connected to the first layer).

[0115] Step 303: According to the dependency matrix DM ij , which divides the elements of each layer in the visual model into multiple groups.

[0116] In some embodiments, according to the dependency matrix DM ij and the elements of each layer in the visual model, perform at least one round of grouping cycle; in each round of grouping cycle, determine any unmarked element in the visual model as the target element corresponding to this round of grouping cycle, and perform at least one round of iterative cycle based on the target element corresponding to this round of grouping cycle; in each round of iterative cycle, mark the first element corresponding to this round of iterative cycle; wherein, the first element corresponding to the first round of iterative cycle is the target element corresponding to this round of grouping cycle; according to the dependency matrix DM ij , determine whether there is a second element in the unmarked elements that has a dependency relationship with the first element corresponding to this round of iterative loop; when there is a second element in the unmarked elements, determine the second element as the first element corresponding to the next round of iterative loop; when there is no second element in the unmarked elements, stop the iterative loop; based on the first element corresponding to each round of iterative loop, obtain a group corresponding to this round of grouping loop; when all elements in each layer of the visual model have been marked, stop the grouping loop.

[0117] As an example, the dependency matrix DM can be constructed based on ij, find all groups as follows:

[0118] 1. Given the i-th layer input or output, first mark the i-th layer input or output, then find all layer inputs and outputs that have dependencies on the i-th layer input or output. The found layer inputs and outputs are marked. Starting from the new layer inputs and outputs found, continue to find new unmarked dependent layer inputs and outputs. The newly found layer inputs and outputs are also marked. This cycle continues until no new dependencies are found. All found layer inputs and outputs form a group with the i-th layer input or output.

[0119] 2. Continue to iterate the remaining unlabeled layer inputs and layer outputs as described in 1 until all layer inputs and layer outputs are labeled.

[0120] Specifically, for Figure 2 The visual model shown, first, for the first layer input Label the first layer input Find and All layer inputs and layer outputs with dependencies, get the output of layer 1 Label layer 1 output Find and All layer inputs and layer outputs with dependencies get the second layer input Label the second layer input Find and All layer inputs and layer outputs with dependencies, none.

[0121] Next, the first layer input Layer 1 output Second layer input All marked, for layer 2 output Label layer 2 output Find and All layer inputs and layer outputs with dependencies get the third layer input Labeling layer 3 inputs Find and All layer inputs and layer outputs with dependencies, none.

[0122] Next, the first layer input Layer 1 output Second layer input Layer 2 output Layer 3 input All marked, for layer 3 output Label layer 3 output Find and All layer inputs and layer outputs with dependencies, none.

[0123] All layer inputs and outputs are labeled and grouped.

[0124] Step 304: Iteratively train the visual model using the image sample training set, and in each round of iterative training, determine the importance index value of the intra-group channel of at least one group based on the weights, weight gradients and attention weights of the elements in at least one group and the corresponding intra-group channels, and update the model parameters of the visual model based on the importance index value of the intra-group channel of at least one group.

[0125] As an example, for Figure 2 The visual model shown has two target groups, namely Target grouping Each layer input and output has 3 channels, and there are 3 channels within the group; for target grouping Each layer input and layer output has 2 channels, and there are 2 channels within a group.

[0126] Figure 2 In the figure, the first column of circles represents the first layer input The second column of circles represents the output of the first layer Channel and second layer input Channel (the second layer is connected to the first layer, the second layer input The channel and the first layer output The channels are the same channels), the third column of circles represents the second layer output Channels and third layer input Channel (the third layer is connected to the second layer, the second layer output The channel and the third layer input The channels are the same channels), the fourth circle represents the output of the third layer channel.

[0127] For the first layer input Assume that the first column of circles (the first layer input The first circle from top to bottom is channel 1, the second circle is channel 2, and the third circle is channel 3; for the first layer output Assume that the second column of circles (the first layer output The first circle from top to bottom is channel 4, the second circle is channel 5, and the third circle is channel 6. Since the first layer is a ReLU layer, or a BN layer, or an Add layer, the input and output channels of the first layer are one-to-one corresponding. Assume that the first column of circles (the first layer input The first circle and the second circle from top to bottom (the first layer output The first circle from top to bottom corresponds one to one; the first column of circles (the first layer input The second circle from top to bottom and the second column of circles (the first layer output The second circle from top to bottom in the channel) corresponds one to one; the first circle (the first layer input The third circle from top to bottom and the second circle (the first layer output The third circle from top to bottom corresponds one to one, that is, channel 1 corresponds one to one with channel 3, channel 2 corresponds one to one with channel 4, and channel 3 corresponds one to one with channel 6. When the importance index value of the intra-group channel is calculated, channel 1 and channel 3 are considered to be one intra-group channel, both of which are intra-group channel 1; channel 2 and channel 4 are considered to be one intra-group channel, both of which are intra-group channel 2; channel 3 and channel 6 are considered to be one intra-group channel, both of which are intra-group channel 3.

[0128] Among them, for the target group The importance index value of the channel within the group is calculated as follows:

[0129] First, the importance index value of any element in the target group to the channel k in the group in any round of iterative training is Imp(k,f), and the importance index value of any element in the target group to the channel k in any round of iterative training is Imp(k,f)=α*Weight fk +β*Grad fk +γ*Att fk , then the first layer input The importance of channel 1 within the group in any round of iterative training

[0130] in, In any round of iterative training, the layer input weights represented by the three lines connected to the first circle (first column) are added together and then weighted. The specific process of weighting is to first calculate in, Among the 9 layer input weights in any round of iterative training, the layer input weights represented by the 3 lines connected to the second circle (first column) are added together. In any round of iterative training, the layer input weights represented by the three lines connected to the third circle (first column) are added together to obtain The weighted process is to input the first layer In any round of iterative training, the weight of channel 1 in the group and the first layer input Regarding target grouping in any round of iterative training The ratio between the sum of the weights of the channels in each group is determined as the first layer input The weight of channel 1 in the group in any round of iterative training.

[0131] In any round of iterative training, the weight gradients corresponding to the 9 layer input weights are added together with the weight gradients corresponding to the layer input weights represented by the 3 lines connected to the first circle (first column), and then weighted. The specific process of weighting is to first calculate in, In any round of iterative training, among the weight gradients corresponding to the 9 layer input weights, the weight gradients corresponding to the layer input weights represented by the 3 lines connected to the second circle (first column) are added together. In any round of iterative training, the weight gradients corresponding to the 9 layer input weights are added together with the weight gradients corresponding to the layer input weights represented by the 3 lines connected to the third circle (first column), and then the weight gradients are obtained. The weighted process is to input the first layer In any round of iterative training, the weight gradient of channel 1 in the group and the first layer input Regarding target grouping in any round of iterative training The ratio of the sum of the weight gradients of each channel in the group is determined as the first layer input The weight gradient of channel 1 in the group in any round of iterative training.

[0132] Then weighting is performed. The specific process of weighting is to first calculate in, Get it again The weighted process is to input the first layer In any round of iterative training, the attention weight of channel 1 in the group and the first layer input Grouping targets in any round of training iteration The ratio of the sum of the attention weights of each channel in the group is determined as the first layer input The weighted attention weight for channel 1 within the group in any round of iterative training.

[0133] First layer output The importance of channel 1 within the group in any round of iterative training

[0134] in, In any round of iterative training, the layer output weights represented by the three lines connected to the first circle (second column) (the left half of the first circle) are added together and then weighted. The specific process of weighting is to first calculate in, Among the 9 layer output weights in any round of iterative training, the layer output weights represented by the 3 lines connected to the second circle (second column) (the left half of the second circle) are added together. In any round of iterative training, the layer output weights represented by the three lines connected to the third circle (second column) (the left half of the third circle) are added together to obtain The weighted process is to output the first layer In any round of iterative training, the weight of channel 1 in the group and the output of the first layer Regarding target grouping in any round of iterative training The ratio between the sum of the weights of the channels in each group is determined as the output of the first layer The weight of channel 1 in the group in any round of iterative training.

[0135] In any round of iterative training, the weight gradients corresponding to the output weights of the 9 layers are added together by the three lines connected to the first circle (the second column) (the left half of the first circle), and then weighted. The specific process of weighting is to first calculate in, In any round of iterative training, among the weight gradients corresponding to the output weights of the nine layers, the weight gradients corresponding to the output weights of the layers represented by the three lines connected to the second circle (second column) (the left half of the second circle) are added together. In any round of iterative training, the weight gradients corresponding to the output weights of the 9 layers are added together by adding the weight gradients corresponding to the output weights of the layers represented by the 3 lines connected to the third circle (second column) (the left half of the third circle), and then we get The weighted process is to output the first layer In any round of iterative training, the weight gradient of channel 1 in the group and the output of the first layer Regarding target grouping in any round of iterative training The ratio of the sum of the weight gradients of each channel in the group is determined as the output of the first layer The weight gradient of channel 1 in the group in any round of iterative training.

[0136] Then weighting is performed. The specific process of weighting is to first calculate in, Get it again The weighted process is to output the first layer In any round of iterative training, the attention weight of channel 1 in the group and the first layer output Grouping targets in any round of training iteration The ratio of the sum of the attention weights of the channels in each group is determined as the output of the first layer The weighted attention weight for channel 1 within the group in any round of iterative training.

[0137] Second layer input The importance of channel 1 within the group in any round of iterative training

[0138] in, In any round of iterative training, the layer input weights represented by the two lines connected to the first circle (second column) (the right half of the first circle) are added together and then weighted. The specific process of weighting is to first calculate in, In any round of iterative training, the layer input weights represented by the two lines connected to the second circle (second column) (the right half of the second circle) are added together. In any round of iterative training, the layer input weights represented by the two lines connected to the third circle (second column) (the right half of the third circle) are added together to obtain The weighted process is to input the second layer In any round of iterative training, the weight of channel 1 in the group and the second layer input Regarding target grouping in any round of iterative training The ratio between the sum of the weights of the channels in each group is determined as the second layer input The weight of channel 1 in the group in any round of iterative training.

[0139] In any round of iterative training, the weight gradients corresponding to the 6 layer input weights are added together with the weight gradients corresponding to the layer input weights represented by the two lines connected to the first circle (second column), and then weighted. The specific process of weighting is to first calculate in, In any round of iterative training, the weight gradients corresponding to the 6 layer input weights are added together, represented by the two lines connected to the second circle (second column). In any round of iterative training, the weight gradients corresponding to the 6 layer input weights are added together with the weight gradients corresponding to the layer input weights represented by the two lines connected to the third circle (second column), and then the weight gradients are obtained. The weighted process is to input the second layer In any round of iterative training, the weight gradient of channel 1 in the group and the second layer input Regarding target grouping in any round of iterative training The ratio of the sum of the weight gradients of each channel in the group is determined as the second layer input The weight gradient of channel 1 in the group in any round of iterative training.

[0140] Then weighting is performed. The specific process of weighting is to first calculate in, Get it again The weighted process is to input the second layer In any round of iterative training, the attention weight of channel 1 in the group and the second layer input Grouping targets in any round of training iteration The ratio of the sum of the attention weights of the channels in each group is determined as the second layer input The weighted attention weight for channel 1 within the group in any round of iterative training.

[0141] In summary,

[0142] The calculation process of intra-group channel 2 and intra-group channel 3 is similar and will not be described here.

[0143] Step 305 : Use the trained visual model to detect the target image to be detected in the coal mine video monitoring system to obtain a detection result of the target image.

[0144] The detection method based on the visual model of the embodiment of the present invention constructs the dependency matrix DM of the visual model according to rules 1, 2 and 3 based on the structure of the visual model. ij ; According to the dependency matrix DM ij, the elements of each layer in the visual model are divided into multiple groups; wherein any element is a layer input or layer output; based on the visual model and the multiple groups, at least one round of iterative training is performed; in each round of iterative training, for any target group including multiple elements, the importance index value of each intra-group channel of the target group in this round of iterative training is determined based on the weight, weight gradient, and attention weight of each element in the target group to the corresponding intra-group channel in this round of iterative training; wherein the importance index value is used to indicate the importance of the corresponding intra-group channel in the target group; according to the importance index value of each intra-group channel of each target group in this round of iterative training, the visual model corresponding to the current round of iterative training is pruned to obtain the visual model corresponding to the next round of iterative training; wherein the visual model corresponding to the first round of iterative training is the visual model; according to the visual model corresponding to the last round of iterative training, the pruned target visual model is obtained. In this way, the input and output of all layers of the model can be grouped based on the layer input-output structure association, thereby realizing the importance evaluation of the model channel group, which can adapt to different model structures and achieve a more refined pruning strategy.

[0145] Let’s take an example to illustrate this. Figure 4 A schematic diagram of a visual model training process provided by an embodiment of the present invention.

[0146] like Figure 4 As shown, the visual model training process may include the following steps:

[0147] Step 1: According to the input-output correlation of different layers in the visual model structure, the layer inputs and layer outputs of all layers of the visual model are divided into several groups (g1, g2, ..., g n ).

[0148] Each group includes mutually dependent layer inputs or layer outputs (e.g. is the first layer input, is the first layer output, is the second layer input).

[0149] Step 2: Add a channel attention module under the convolutional layer of the visual model.

[0150] Step 3: Enter the number of training times Iter Pruned , each iteration updates the ratio Ratio Pruned , model weight mask Mask = 1, image sample training set D.

[0151] Step 4: Train the model, save the model weight gradient, and update the model weight; if the mask value of the corresponding weight is 0, the gradient of the weight is not updated.

[0152] The visual model is trained based on the image sample training set D, all the weight gradients of the model are saved, and the model weights are updated. If the Mask value of the corresponding weight is 0, the gradient is set to 0 and the weight is not updated.

[0153] Step 5: For each group g i As a unit, calculate the importance Imp(k) of each channel in the group.

[0154] Step 6: For each group g i The unit is c_g, and the number of channels in the group is c_g i , ignore the channels with Mask 0 in the group (that is, the channel corresponding to the layer input weight or layer output weight Mask in the group is 0), sort the remaining channels from small to large according to their importance Imp(k), and select the top Ratio Pruned *c_g i Channels and set all layer input weight Masks and layer output weight Masks in the group corresponding to these channels to 0.

[0155] Each group g i The unit is c_g, and the number of channels in the group is c_g i , ignore the channels with Mask 0 in the group (that is, the channel corresponding to the layer input weight or layer output weight Mask in the group is 0), sort the remaining channels from small to large according to their importance Imp(k), and select the top Ratio Pruned *c_g i Channels, and set all layer input weights and masks as well as layer output weights and masks in the group corresponding to these channels to 0.

[0156] Step 6: Determine whether to end the training? (Is the current training number less than the set training number Iter Pruned )

[0157] Determine whether the current training times are lower than the set training times. If the current training times are equal to the set training times, end the training and proceed to the next step. Otherwise, jump to step 4.

[0158] Step 7: Remove the channel attention module added in step 2 and the weight with Mask equal to 0 to obtain the trained visual model.

[0159] Step 8: Fine-tune the trained visual model.

[0160] The above visual model training process adopts the dynamic group channel pruning method based on structural dependence proposed by the present invention. This method is a technique that systematically prunes redundant or unimportant channels by analyzing and utilizing the dependence relationships between layers and within channels in a neural network, thereby reducing model parameters and computational volume while maintaining the model performance as much as possible. This method mainly includes two parts: the model pruning framework for dynamic group channels based on structural dependence, and the channel importance evaluation method based on the structural association between layer inputs and outputs. Among them, the model pruning framework for dynamic group channels based on structural dependence is the model pruning framework described in steps 1 to 8 above. The content of the channel importance evaluation method based on the structural association between layer inputs and outputs is as follows:

[0161] 1. Group the inputs and outputs of all layers of the model based on the structural association between layer inputs and outputs. The specific workflow is as follows:

[0162] A layer is the basic building block of a deep learning model. Each layer generates an output from the input through specific mathematical operations. Multiple layers are combined to form a complex model. The inputs and outputs of each layer are respectively defined as Assume the model has L layers. All layer inputs in the model All layer outputs

[0163] Initialize an input-output layer dependence matrix DM = 0 2L×2L , where the meaning of DM in the matrix ij is as follows:

[0164]

[0165] DM ij = 0 indicates no dependence, and DM ij = 1 indicates dependence;

[0166] Assume that the i-th layer and the j-th layer in the model structure are connected, i < j. Then DM (i+L)j = 1, and DM j(i+L) = 1;

[0167] Assume that the input and output channels of the i-th layer in the model structure are in one-to-one correspondence (such as relu layer, bn layer, add layer). Then DM (i+L)i = 1, and DM i(i+L) = 1;

[0168] The connection is transitive. If DM (i+L)j = 1 and DM j(k+L) = 1, then DM (i+L)(k+L) = 1.

[0169] According to the constructed input-output layer dependence matrix, find all groups according to the following steps:

[0170] 1. Given the i-th layer input or output, first mark the i-th layer input or output, then find all layer inputs and outputs that have dependencies on the i-th layer input or output. The found layer inputs and outputs are marked. Starting from the new layer inputs and outputs found, continue to find new unmarked dependent layer inputs and outputs. The newly found layer inputs and outputs are also marked. This cycle continues until no new dependencies are found. All found layer inputs and outputs form a group with the i-th layer input or output.

[0171] 2. Continue to iterate the remaining unlabeled layer inputs and layer outputs as described in 1 until all layer inputs and layer outputs are labeled.

[0172] 2. The evaluation method for the importance of channels within a group is based on three indicators: weight size, weight gradient size, and channel attention output weight. The specific workflow is as follows:

[0173] Group g i Contains multiple layer inputs (such as ) and layer outputs (such as ), each layer input and output has K channels. The importance of layer input or layer output to channel k is defined as Imp(k,f). The importance of channel k within a group is obtained by summing the importance of all layer inputs and layer outputs to channel k: Imp(k) = ∑ f Imp(k,f).

[0174] The importance of a layer input or layer output f to the kth channel is obtained by weighted summing the layer weight, weight gradient, and channel attention output weight:

[0175] Imp(k,f)=α*Weight fk +β*Grad fk +γ*Att fk

[0176] The weight is calculated as follows:

[0177]

[0178] The weight gradient size is calculated as follows:

[0179]

[0180] The channel attention output weight is calculated as follows:

[0181]

[0182] Among them, the attention weight is determined by the output of the sigmoid layer in the channel attention module, and the channel attention module includes a global average pooling layer, two fully connected layers and a sigmoid layer.

[0183] Compared to related technologies, structured pruning methods often lack versatility across diverse model structures, while unstructured pruning, while easy to implement, often significantly impacts model performance. The aforementioned dynamic group channel pruning method, based on structural dependencies, constructs a channel structure dependency graph to perform structured pruning for models of specific scale and structure, while maintaining model accuracy as much as possible. This method is adaptive to diverse model structures, enabling more refined pruning strategies.

[0184] Compared to related technologies, channel importance assessment methods often focus on evaluating the impact of channels on the overall model output or predictions of specific categories, lacking a comprehensive and systematic evaluation of channel importance. This method, based on the correlation between layer input and output structures, integrates multiple evaluation metrics to identify the channels that contribute least to model performance. This method employs a multi-level evaluation strategy to ensure that pruned channels have a minimal impact on the model output, thereby reducing model parameters and computational complexity while maintaining the model's predictive accuracy.

[0185] In summary, this paper implements an adaptive structured pruning strategy for model structures through a dynamic group channel pruning method based on structural dependencies. This pruning method is applicable to a variety of deep learning model structures and eliminates the need to manually establish dependencies between model structures, simplifying the model pruning process and saving labor costs. While maintaining model accuracy, this method reduces the model's computational complexity and storage requirements, making it potentially applicable in resource-constrained environments.

[0186] Furthermore, through a channel importance assessment method based on the correlation between layer input and output structures, the importance of model channel groups is assessed. By integrating multiple importance assessment techniques, the impact of each channel group on model performance is identified and ranked, thereby selecting the channels with the lowest impact on model accuracy for pruning, ensuring that the model's performance loss after compression is minimized. This method dynamically adjusts the pruning intensity to minimize information loss during the pruning process, improving compression efficiency and model performance, and providing a reliable and economical solution for the efficient deployment and application of deep learning models.

[0187] In order to implement the above embodiment, the present invention also proposes a detection device based on a visual model.

[0188] Figure 5 A schematic structural diagram of a detection device based on a visual model provided in an embodiment of the present invention.

[0189] like Figure 5 As shown, the detection device based on the visual model includes: an acquisition module 51, a grouping module 52, a training module 53 and a detection module 54.

[0190] An acquisition module 51 is configured to acquire an image sample training set of any visual model in a coal mine video monitoring system, wherein the image sample training set includes a detection object and detection information of the detection object, and the visual model is a network model;

[0191] a grouping module 52 for dividing the elements into a plurality of groups according to the dependency relationship between the elements of each layer in the visual model, wherein the elements are layer inputs or layer outputs of each layer;

[0192] A training module 53 is configured to iteratively train the visual model using the image sample training set, and in each round of iterative training, determine an importance index value of a channel within the at least one group based on the weights, weight gradients, and attention weights of the elements in the at least one group and the corresponding channel within the group, and update model parameters of the visual model based on the importance index value of the channel within the at least one group, wherein the importance index value is used to indicate the importance of the channel within the group in the corresponding group;

[0193] The detection module 54 is used to detect the target image to be detected in the coal mine video monitoring system using the trained visual model to obtain the detection result of the target image.

[0194] Furthermore, in a possible implementation of the embodiment of the present invention, the training module 53 includes:

[0195] An acquisition unit is used to iteratively train the visual model corresponding to the current round of iterative training using the image sample training set in each round of iterative training, and obtain the weight and weight gradient of each layer element in the visual model in the current round of iterative training; wherein the visual model corresponding to the first round of iterative training is the visual model;

[0196] A determination unit is configured to determine, for any target group including multiple elements, an importance index value of each intra-group channel of the target group in this round of iterative training based on the weight and weight gradient of each element in the target group in this round of iterative training, and the attention weight of each element in the target group to the corresponding intra-group channel in this round of iterative training;

[0197] The updating unit is used to update the model parameters of the visual model corresponding to the current round of iterative training according to the importance index value of each channel in each group of each target group in the current round of iterative training, so as to obtain the visual model corresponding to the next round of iterative training.

[0198] Furthermore, in a possible implementation of the embodiment of the present invention, the updating unit is further configured to:

[0199] For any target group, determine candidate channels from the channels within the target group based on the correspondence between each element in the target group and the channels within the target group and the weight mask values ​​of each element in the target group in this round of iterative training; wherein the weight mask values ​​of the elements corresponding to the candidate channels in this round of iterative training are not the first set values, and the weight mask values ​​of the elements of each layer in the visual model in the first round of iterative training are all the second set values;

[0200] Sort the candidate channels according to their importance index values ​​to select a target number of target channels from the sorted candidate channels; the target number is determined based on the iterative update ratio corresponding to the current round of iterative training and the total number of channels in the target group;

[0201] The weights and weight mask values ​​of the elements corresponding to the selected target number of target channels are set to the first set values ​​to obtain the visual model corresponding to the next round of iterative training.

[0202] Furthermore, in a possible implementation of an embodiment of the present invention, in any round of iterative training, when the weight mask value of any element in the previous round of iterative training is the first set value, the weight gradient of the element in this round of iterative training is set to the first set value, and no update operation is performed on the weight of the element in this round of iterative training.

[0203] Furthermore, in a possible implementation of the embodiment of the present invention, the determining unit is further configured to:

[0204] For any target group consisting of multiple elements, determine the importance index value of the element to the target channel in this round of iterative training based on the weight and weight gradient of any element in this round of iterative training, as well as the element's attention weight to the corresponding intra-group channel in this round of iterative training; where the target channel is any intra-group channel corresponding to the element;

[0205] The importance index value of each element in the target group to the target channel in this round of iterative training is accumulated to obtain the importance index value of the target channel in this round of iterative training.

[0206] Furthermore, in a possible implementation of the embodiment of the present invention, the determining unit is further configured to:

[0207] For any target group including multiple elements, determine the weight of each channel within the target group in the current round of iterative training according to the weight of the target element in the current round of iterative training; wherein the target element is any element in the target group;

[0208] Determine the ratio between the weight of the target element with respect to the target channel in the current iteration training and the sum of the weights of the target element with respect to each intra-group channel of the target group in the current iteration training as the weight of the target element with respect to the target channel in the current iteration training;

[0209] Determine the weight gradient of the target element with respect to each intra-group channel of the target group in the current iteration training according to the weight gradient of the target element in the current iteration training;

[0210] Determine the ratio between the weight gradient of the target element with respect to the target channel in the current iteration training and the sum of the weight gradients of the target element with respect to each intra-group channel of the target group in the current iteration training as the weight gradient of the target element with respect to the target channel in the current iteration training;

[0211] Determine the ratio between the attention weight value of the target element with respect to the target channel in the current iteration training and the sum of the attention weight values of the target element with respect to each intra-group channel of the target group in the current iteration training as the weighted attention weight value of the target element with respect to the target channel in the current iteration training;

[0212] Based on the weight of the target element with respect to the target channel in the current iteration training, the weight gradient of the target element with respect to the target channel in the current iteration training, and the weighted attention weight value of the target element with respect to the target channel in the current iteration training, determine the importance index value of the target element with respect to the target channel in the current iteration training.

[0213] Further, in a possible implementation manner of the embodiment of the present invention, the grouping module 52 includes:

[0214] A construction unit for constructing a dependency matrix DM of the visual model based on Rules 1, Rules 2, and Rules 3 according to the structure of the visual model ij ; where

[0215]

[0216] DM ij =0 indicates no dependency relationship, DM ij =1 indicates a dependency relationship, L represents the number of layers of the visual model, 0 < i ≤ 2L, 0 < j ≤ 2L,

[0217] The dependency matrix DM ij is used to indicate the dependency relationship between the elements of each layer in the visual model,

[0218] Rule 1 is that if the i-th layer and the j-th layer in the structure of the visual model are connected, i < j, then DM (i+L)j =1, DM j(i+L) =1,

[0219] Rule 2 states that if the input and output channels of the i-th layer in the structure of the visual model are one-to-one corresponding, then DM (i+L)i =1,DM i(i+L) =1,

[0220] Rule 3 states that a connection is transitive if the DM (i+L)j =1 and DM j(k+L) =1, then DM (i+L)(k+L) =1;

[0221] The grouping unit is used to group the ij , which divides the elements of each layer in the visual model into multiple groups.

[0222] Furthermore, in a possible implementation of the embodiment of the present invention, the grouping unit is further configured to:

[0223] According to the dependency matrix DM ij and elements of each layer in the visual model, performing at least one round of grouping loop;

[0224] In each grouping cycle, any unlabeled element in the visual model is determined as the target element corresponding to this grouping cycle, and at least one iterative cycle is performed based on the target element corresponding to this grouping cycle;

[0225] In each round of iteration, the first element corresponding to the current round of iteration is marked; wherein the first element corresponding to the first round of iteration is the target element corresponding to the current round of grouping;

[0226] According to the dependency matrix DM ij , determine whether there is a second element in the unmarked elements that has a dependency relationship with the first element corresponding to this round of iteration loop;

[0227] If there is a second element in the unmarked elements, the second element is determined as the first element corresponding to the next iteration cycle;

[0228] If the second element does not exist in the unmarked elements, stop the iteration loop;

[0229] Based on the first element corresponding to each round of iterative loop, a group corresponding to this round of grouping loop is obtained;

[0230] When all elements in each layer of the visual model have been marked, stop the grouping loop.

[0231] Furthermore, in a possible implementation of the embodiment of the present invention, the attention weight is determined based on an output value or a preset value of a channel attention module added under the convolutional layer of the visual model; the above-mentioned device also includes:

[0232] A determination module is used to determine whether to stop iterative training according to the number of iterative training;

[0233] The processing module is used to remove the weights of the added channel attention module and the visual model corresponding to the last round of iterative training, whose weight mask value is the first set value, to obtain a trained visual model.

[0234] It should be noted that the above explanation of the embodiment of the detection method based on the visual model is also applicable to the detection device based on the visual model of this embodiment, and will not be repeated here.

[0235] The visual model-based detection device of an embodiment of the present invention obtains an image sample training set of any visual model in a coal mine video surveillance system; divides each element into multiple groups based on the dependency relationship between each layer element in the visual model, wherein the element is a layer input or layer output of each layer; iteratively trains the visual model using the image sample training set, and in each round of iterative training, determines the importance index value of the intra-group channel of at least one group based on the weight, weight gradient and attention weight of the corresponding intra-group channel of the element in at least one group, and updates the model parameters of the visual model based on the importance index value of the intra-group channel of at least one group; uses the trained visual model to detect the target image to be detected in the coal mine video surveillance system to obtain the detection result of the target image, which can compress the visual model structure while maintaining the detection accuracy. Therefore, by dividing the groups based on the dependency relationship between the layer input and layer output of each layer in the visual model, and determining the importance index value of the intra-group channel of at least one group, unimportant channels are pruned based on the importance index value of the intra-group channel of at least one group, thereby reducing model parameters and computational complexity, and achieving universality while maintaining the model accuracy as much as possible.

[0236] In order to implement the above embodiments, the present invention also proposes an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the visual model-based detection method proposed in any of the above embodiments of the present invention.

[0237] Figure 6 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. It should be noted that: Figure 6 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present invention.

[0238] like Figure 6As shown, the electronic device may include: a shell 11, a processor 12, a memory 13, a circuit board 14 and a power supply circuit 15, wherein the circuit board 14 is placed inside the space enclosed by the shell 11, and the processor 12 and the memory 13 are arranged on the circuit board 14; the power supply circuit 15 is used to supply power to various circuits or devices of the above-mentioned electronic device; the memory 13 is used to store executable program code; the processor 12 runs the program corresponding to the executable program code by reading the executable program code stored in the memory 13, so as to execute the visual model-based detection method proposed in any of the above embodiments of the present invention.

[0239] For details on the specific execution process of the above steps by the processor 12 and the steps further executed by the processor 12 by running the executable program code, please refer to the present invention. Figure 1 The description of the illustrated embodiment will not be repeated here.

[0240] In order to implement the above embodiments, the present invention further proposes a computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the visual model-based detection method proposed in any of the above embodiments of the present invention.

[0241] In order to implement the above embodiments, the present invention further proposes a computer program product, including a computer program, which, when executed by a processor, implements the visual model-based detection method proposed in any of the above embodiments of the present invention.

[0242] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0243] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0244] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.

[0245] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.

[0246] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0247] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0248] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium.

[0249] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and are not to be construed as limiting the present invention. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A detection method based on a visual model, characterized in that: include: Obtain an image sample training set of any visual model in a coal mine video monitoring system, wherein the image sample training set includes a detection object and detection information of the detection object, and the visual model is a network model; Dividing the elements of each layer in the visual model into a plurality of groups according to dependencies between the elements, wherein the elements are layer inputs or layer outputs of each layer; Iteratively training the visual model using the image sample training set, and in each round of iterative training, determining an importance index value of a channel within at least one of the groups based on a weight, a weight gradient, and an attention weight of an element in the group, and updating model parameters of the visual model based on the importance index value of the channel within the group, wherein the importance index value is used to indicate the importance of the channel within the group in the corresponding group; The trained visual model is used to detect the target image to be detected in the coal mine video monitoring system to obtain the detection result of the target image.

2. The method according to claim 1, characterized in that The iterative training of the visual model using the image sample training set, and in each round of iterative training, determining an importance index value of a channel within at least one of the groups according to a weight, a weight gradient, and an attention weight of an element in at least one of the groups and a corresponding channel within the group, and updating a model parameter of the visual model according to the importance index value of the channel within the group, includes: In each round of iterative training, the visual model corresponding to the current round of iterative training is iteratively trained using the image sample training set to obtain the weights and weight gradients of each layer element in the visual model in the current round of iterative training; wherein the visual model corresponding to the first round of iterative training is the visual model; For any target group including multiple elements, the importance index value of each intra-group channel of the target group in this round of iterative training is determined according to the weight and weight gradient of each element in the target group in this round of iterative training, and the attention weight of each element in the target group to the corresponding intra-group channel in this round of iterative training; according to the importance index value of each intra-group channel of each target group in this round of iterative training, the model parameters of the visual model corresponding to this round of iterative training are updated to obtain the visual model corresponding to the next round of iterative training.

3. The method according to claim 2, characterized in that The updating of the model parameters of the visual model corresponding to the current round of iterative training according to the importance index value of each channel within each target group in the current round of iterative training to obtain the visual model corresponding to the next round of iterative training includes: For any of the target groups, determining candidate channels from the intra-group channels of the target group according to the correspondence between each of the elements in the target group and the intra-group channels of the target group and the weight mask value of each of the elements in the target group in this round of iterative training; wherein the weight mask values ​​of the elements corresponding to the candidate channels in this round of iterative training are not the first set values, and the weight mask values ​​of the elements of each layer in the visual model in the first round of iterative training are all the second set values; Sorting the candidate channels according to the importance index values ​​of the candidate channels to select a target number of target channels from the sorted candidate channels; wherein the target number is determined based on the iterative update ratio corresponding to the current round of iterative training and the total number of channels in the target group; The weights and weight mask values ​​of the elements corresponding to the selected target number of target channels are set to the first set values ​​to obtain the visual model corresponding to the next round of iterative training.

4. The method according to claim 2, characterized in that In any round of iterative training, if the weight mask value of any element in the previous round of iterative training is the first set value, the weight gradient of the element in this round of iterative training is set to the first set value, and no update operation is performed on the weight of the element in this round of iterative training.

5. The method according to claim 2, characterized in that For any target group including a plurality of elements, determining the importance index value of each intra-group channel of the target group in this round of iterative training according to the weight and weight gradient of each element in the target group in this round of iterative training, and the attention weight of each element in the target group to the corresponding intra-group channel in this round of iterative training, including: For any target group including multiple elements, determine the importance index value of the element to the target channel in this round of iterative training based on the weight and weight gradient of any element in this round of iterative training, and the attention weight of the element to the corresponding intra-group channel in this round of iterative training; wherein the target channel is any intra-group channel corresponding to the element; The importance index values ​​of the elements in the target group on the target channel in this round of iterative training are accumulated to obtain the importance index value of the target channel in this round of iterative training.

6. The method according to claim 5, characterized in that For any target group including a plurality of the elements, determining the importance index value of the element to the target channel in this round of iterative training according to the weight and weight gradient of any element in this round of iterative training, and the attention weight of the element to the corresponding channel within the group in this round of iterative training, includes: For any target group including a plurality of the elements, determining the weight of each channel within the target group in the current round of iterative training for the target element according to the weight of the target element in the current round of iterative training; wherein the target element is any element in the target group; Determine the weight of the target element on the target channel in this round of iterative training as the weight of the target element on the target channel in this round of iterative training; Determining, according to the weight gradient of the target element in this round of iterative training, the weight gradient of each intra-group channel of the target group of the target element in this round of iterative training; Determine the ratio of the weight gradient of the target element with respect to the target channel in this round of iterative training to the sum of the weight gradients of the target element with respect to each channel within the target group in this round of iterative training as the weight gradient of the target element with respect to the target channel in this round of iterative training; Determine the weighted attention weight of the target element to the target channel in this round of iterative training as the ratio of the attention weight of the target element to the target channel in this round of iterative training to the sum of the attention weights of the target element to the channels in each group of the target group in this round of iterative training; Based on the weight of the target element to the target channel in this round of iterative training, the weight gradient of the target element to the target channel in this round of iterative training and the weighted attention weight of the target element to the target channel in this round of iterative training, the importance index value of the target element to the target channel in this round of iterative training is determined.

7. The method according to claim 1, characterized in that The step of dividing the elements into a plurality of groups according to the dependency relationship between the elements of each layer in the visual model comprises: According to rules 1, 2 and 3, based on the structure of the visual model, the dependency matrix DM of the visual model is constructed. ij ;in, DM ij =0 means no dependency, DM ij =1 indicates a dependency relationship, L indicates the number of layers of the visual model, 0<i≤2L, 0<j≤2L, The dependency matrix DM ij used to indicate the dependency relationship between the elements of each layer in the visual model, The rule 1 is that if the i-th layer and the j-th layer in the structure of the vision model are connected, where i < j, then DM (i+L)j = 1, DM j(i+L) = 1, Rule 2 states that if the input and output channels of the i-th layer in the structure of the visual model are one-to-one corresponding, then DM (i+L)i =1,DM i(i+L) =1, Rule 3 states that a connection is transferable if the DM (i+L)j =1 and DM j(k+L) =1, then DM (i+L)(k+L) =1; According to the dependency matrix DM ij , dividing the elements of each layer in the visual model into multiple groups.

8. The method according to claim 7, characterized in that According to the dependency matrix DM ij , dividing the elements of each layer in the visual model into a plurality of groups, including: According to the dependency matrix DM ij and the elements of each layer in the visual model, performing at least one grouping cycle; In each grouping cycle, any unlabeled element in the visual model is determined as a target element corresponding to the current grouping cycle, and at least one iterative cycle is performed based on the target element corresponding to the current grouping cycle; In each round of iteration, the first element corresponding to the current round of iteration is marked; wherein the first element corresponding to the first round of iteration is the target element corresponding to the current round of grouping; According to the dependency matrix DM ij , determining whether there is a second element in the unmarked elements that has a dependency relationship with the first element corresponding to the current iteration loop; If the second element exists in the unmarked elements, determining the second element as the first element corresponding to the next round of iteration loop; If the second element does not exist in the unmarked elements, stop the iteration loop; Based on the first element corresponding to each round of iterative loop, a group corresponding to this round of grouping loop is obtained; When all elements of the layers in the visual model have been marked, the grouping loop is stopped.

9. The method according to any one of claims 1 to 8, characterized in that The attention weight is determined based on an output value or a preset value of a channel attention module added under the convolutional layer of the visual model; the method further includes: Determine whether to stop iterative training according to the number of iterative training; The added channel attention module and the weight of the visual model corresponding to the last round of iterative training whose weight mask value is the first set value are removed to obtain the trained visual model.

10. A detection device based on a visual model, characterized in that: include: An acquisition module is used to acquire an image sample training set of any visual model in a coal mine video monitoring system, wherein the image sample training set includes a detection object and detection information of the detection object, and the visual model is a network model; a grouping module, configured to divide the elements of each layer in the visual model into a plurality of groups according to dependencies between the elements of each layer, wherein the elements are layer inputs or layer outputs of each layer; a training module, configured to iteratively train the visual model using the image sample training set, and in each round of iterative training, determine an importance index value of a channel within at least one of the groups based on the weights, weight gradients, and attention weights of the elements in at least one of the groups and the corresponding channel within the group, and update model parameters of the visual model based on the importance index value of the channel within the group, wherein the importance index value is used to indicate the importance of the channel within the group in the corresponding group; The detection module is used to detect the target image to be detected in the coal mine video monitoring system using a trained visual model to obtain a detection result of the target image.

Citation Information

Patent Citations

  • Underground coal mine lightweight real-time intelligent video monitoring system for pruning based on attention mechanism

    CN115410127A

  • Pruning method and device of neural network and neural network construction system

    CN117077760A

  • Image detection method based on SAR target detector

    CN117557902A

  • Improved YOLOv7 garbage detection network model lightweight method

    CN118262217A

  • Low-overhead multipoint time-frequency positioning method based on federated learning framework

    CN119277324A

Cited By

  • Mine ai video monitoring detection and identification method based on combination of small model and large model

    CN122761284A