Intelligent identification and disposal method and device for blind shot of blasting based on multi-modal perception

By using multimodal perception fusion technology and deep learning models to identify misfires, the problems of low accuracy of single perception methods and high risk of manual inspection are solved, achieving high-precision detection and intelligent handling of misfires and reducing the risks of mining operations.

CN122391798APending Publication Date: 2026-07-14POWERCHINA RAILWAY CONSTR +3
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
POWERCHINA RAILWAY CONSTR
Filing Date
2026-03-09
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing single-sensor methods have low accuracy in detecting misfires at mine blasting sites, and manual inspection is risky, making it difficult to meet safety requirements in complex environments.

Method used

A multimodal perception method is adopted, which combines visible light images, infrared thermal imaging and lidar point cloud data. Features are extracted and fused through a deep learning model to generate blind shot detection results, and intelligent handling strategies are provided based on the results.

Benefits of technology

It significantly reduces the rate of missed detections and false detections, reduces the risk of manual entry into blasting operation areas, improves detection accuracy and safety, and is in line with the development trend of mine mechanization and automation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122391798A_ABST
    Figure CN122391798A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multi-modal perception's blasting blind cannon intelligent identification and disposal method and device, it is related to mine blasting safety and intelligent monitoring technical field, by collecting multi-modal perception data, and using deep learning to the multi-modal perception data is analyzed, and then determine blind cannon detection result and corresponding disposal strategy, can effectively overcome the limitation of single sensor under complex blasting site, greatly reduce the missed detection rate and false detection rate, staff need not directly enter the blasting operation area full of unexploded blind cannon risk, greatly reduce the casualty risk of operating personnel, comply with the development trend of mine safety mechanization man replacement, automation reduces person.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of mine blasting safety and intelligent monitoring technology, and more specifically, to a method and device for intelligent identification and handling of misfires based on multimodal perception. Background Technology

[0002] In blasting operations such as mining and tunneling, misfires (misfired charges) pose a serious safety hazard. If misfires are not detected and handled properly in a timely manner, they may accidentally detonate during subsequent operations such as loading and drilling, resulting in casualties and equipment damage. Traditional misfire detection mainly relies on manual visual inspection, which has many problems: First, the post-blasting environment is harsh, with dust, smoke, and loose rocks, making manual entry extremely risky; second, misfires are often covered by blast piles or buried deep in the hole, making them difficult to spot with the naked eye; third, manual inspection is highly subjective and easily affected by factors such as fatigue and lighting, leading to missed or false detections.

[0003] In recent years, with the development of machine vision and sensor technology, image recognition-based methods for detecting unexploded ordnance have been gradually proposed. However, existing single-sensor methods often have limitations. For example, visible light images are easily affected by insufficient lighting and dust obstruction, making it difficult to distinguish residual explosives in deep holes; infrared thermal imaging, while capable of detecting temperature anomalies caused by unexploded reactions, is easily interfered with by ambient background temperature and thermal radiation from rocks; lidar can acquire the three-dimensional morphology of the blast pile, but it is difficult to directly distinguish the material properties (such as rocks and explosives) within the blast pile. Therefore, the detection accuracy of single-mode methods is low and cannot meet the actual needs of complex blasting sites. Summary of the Invention

[0004] The present invention provides an intelligent identification and handling method for misfires based on multimodal perception, which aims to solve the problems of low detection accuracy of single perception methods and high risk of manual inspection.

[0005] This invention provides a method for intelligent identification and handling of misfires based on multimodal perception, comprising: After the blasting operation is completed and the prescribed smoke extraction waiting time has elapsed, multimodal sensing data corresponding to the blasting operation area is collected; wherein, the multimodal sensing data includes visible light image sequences, infrared thermal imaging image sequences, and lidar point cloud data; A first deep learning model is used to extract the first feature corresponding to the visible light image sequence, a second deep learning model is used to extract the second feature corresponding to the infrared thermal imaging image sequence, and a third deep learning model is used to extract the third feature corresponding to the lidar point cloud data. The first feature, the second feature, and the third feature are fused to obtain a fused feature, and a fourth deep learning model is used to identify the fused feature to obtain the blind shot detection result; Based on the results of the misfire detection, a corresponding handling strategy is formulated for the misfire detection area, and the handling strategy is fed back to the staff to realize intelligent identification and handling of blasting misfires based on multimodal perception.

[0006] Furthermore, the first and second deep learning models are both set to YOLOv5, CNN, or Faster R-CNN, and the third deep learning model is set to PointNet++.

[0007] Furthermore, before using the first, second, and third deep learning models, a multi-strategy optimization algorithm is employed for training, including: Use the first deep learning model, the second deep learning model, or the third deep learning model as the target deep learning model; The model parameters of the target deep learning model are initialized and encoded into vectors multiple times to obtain multiple individual model parameters; Obtain the fitness of each individual model parameter, and determine the optimal individual among the individual model parameters based on the fitness. The model parameter individuals are subjected to optimal noise sampling on the optimal individual to obtain the model parameter individuals after optimal noise sampling; Based on the dispersion of individual model parameters in the solution space after optimal noise acquisition, local augmentation decision parameters and global augmentation decision parameters are obtained. Decisions are made based on the local enhancement decision parameters, and the individual model parameters after the optimal noise acquisition are selectively and continuously fitted to obtain the individual model parameters after continuous fitting. Decisions are made based on the global enhancement decision parameters, and the individual model parameters after continuous fitting are selectively mutated to obtain the mutated individual model parameters. Determine whether the number of training iterations has reached the preset maximum number of training iterations. If so, determine the final model parameters of the target deep learning model based on the individual model parameters after the mutation process. Otherwise, based on the target deep learning model, return to the step of obtaining the optimal individual.

[0008] Furthermore, the model parameter individuals are subjected to optimal noise sampling on the optimal individual to obtain the model parameter individuals after optimal noise sampling, including: The model parameters are used to collect noise from the optimal individual, and the disturbance information is as follows: ; ; In the formula, For the first i The first model parameter of the individual d The perturbation of the dimensional parameter, i =1,2,…,M, where M is the total number of individual model parameters. d =1,2,…,NP, where NP is the total number of parameters in each individual model parameter. Let be the first constant between (0.1, 3). The first random number between (0,1) For intermediate parameters, For the first i The Euclidean distance between each model parameter individual and the optimal individual Pi The second random number between (0,1) The second constant is set to 1 or 2, t is the number of training iterations, and T is the maximum number of training iterations; Based on the disturbance information and the optimal individual, the optimal model parameter individual after noise acquisition is obtained as follows: ; in, The first optimal individual d Dimensional parameters, For the first i The model parameters of the individual after the optimal noise acquisition. d Dimensional parameters, The third random number between (0,1) The fourth random number between (0,1) As an information guiding factor, and , For the first i +1 model parameter individual's first d dimensional parameters, and i When it is M, For the individual parameters of the random model d Dimensional parameters.

[0009] Furthermore, based on the dispersion of individual model parameters in the solution space after optimal noise acquisition, local augmentation decision parameters and global augmentation decision parameters are obtained, including: The dispersion of individual model parameters in the solution space after obtaining the optimal noise acquisition is as follows: ; In the formula, For dispersion parameters, This is the scaling parameter, and it is set to 80. For the firstk The fitness of individual model parameters after optimal noise acquisition The average fitness of individual model parameters after all optimal noise acquisitions. k =1,2,…,M; Based on the aforementioned dispersion, the local augmentation decision parameters and the global augmentation decision parameters are determined as follows: ; ; In the formula, To locally enhance decision parameters, To enhance decision parameters globally, For the first h The training progress factor for each successive fitting operation. The total number of times continuous fitting is performed. For the first g Training progress factor for each mutation process. The training progress factor is the total number of mutation operations performed. , Fitness is the fitness level prior to performing continuous fitting or mutation treatment. Max is the function that takes the maximum value after performing continuous fitting or mutation processing.

[0010] Further, based on the local enhancement decision parameters, a decision is made, and the individual model parameters after the optimal noise acquisition are selectively and continuously fitted to obtain the individual model parameters after continuous fitting, including: The first decision factor is randomly generated between (0, 1); When the first decision factor is less than the local enhancement decision parameter, the individual model parameters after the optimal noise acquisition are continuously fitted to obtain the individual model parameters after continuous fitting: ; ; ; In the formula, For the first m Individual model parameters after optimal noise acquisition For the first m Individual model parameters after continuous fitting. The first random individual is randomly selected from the model parameters individuals after optimal noise acquisition. This refers to the second random individual selected from the model parameter individuals after optimal noise acquisition. The third random individual is randomly selected from the model parameter individuals after the optimal noise acquisition. It is a natural constant. For the optimal individual, The first fitting coefficient, The second fitting coefficient, for , as well as The corresponding mean individual, for The corresponding advantages and disadvantages of the location, for The corresponding advantages and disadvantages of the location, for The corresponding advantages and disadvantages of the location, The basic weighting coefficient is set to 0.01; Location superiority / inferiority = ; For the first j The fitness of a random individual The fitness of the optimal individual.

[0011] Further, based on the global enhancement decision parameters, a decision is made, and selective mutation processing is performed on the individual model parameters after continuous fitting to obtain mutated individual model parameters, including: A second decision factor is randomly generated between (0, 1); When the second decision factor is less than the global augmentation decision parameter, the individual model parameters after continuous fitting are mutated to obtain the mutated individual model parameters as follows: ; ; In the formula, For the first n The first individual of the model parameters after the nth consecutive fitting d Dimensional parameters, For the first n Individual model parameters after mutation treatment As a variable factor, This is the variation range control factor, and it is set to 0.45; The order of mutation, This represents the total mutation order, and is set to 3. For the first d The upper limit of the dimension parameter.

[0012] Further, a first feature corresponding to the visible light image sequence is extracted using a first deep learning model, a second feature corresponding to the infrared thermal imaging image sequence is extracted using a second deep learning model, and a third feature corresponding to the lidar point cloud data is extracted using a third deep learning model, including: The visible light images in the visible light image sequence are input one by one into the first deep learning model, and the output of the first deep learning model is used to construct the first feature. The infrared thermal imaging images in the infrared thermal imaging image sequence are input one by one into the second deep learning model, and the output of the second deep learning model is used to construct the second feature. The lidar point cloud data is input into the third deep learning model to obtain the third feature corresponding to the lidar point cloud data.

[0013] Furthermore, the disposal strategy includes at least: a grouting failure disposal strategy and a mechanical detonation strategy.

[0014] On the other hand, the present invention provides an intelligent identification and handling device for misfired explosives based on multimodal perception, comprising: The multimodal data acquisition module is used to acquire multimodal sensing data corresponding to the blasting operation area after the blasting operation is completed and a specified smoke extraction waiting time has elapsed; wherein, the multimodal sensing data includes visible light image sequences, infrared thermal imaging image sequences, and lidar point cloud data; The feature extraction module is used to extract the first feature corresponding to the visible light image sequence using a first deep learning model, extract the second feature corresponding to the infrared thermal imaging image sequence using a second deep learning model, and extract the third feature corresponding to the lidar point cloud data using a third deep learning model. The blind shot detection module is used to fuse the first feature, the second feature and the third feature to obtain the fused feature, and to use a fourth deep learning model to identify the fused feature to obtain the blind shot detection result. The handling suggestion module is used to formulate a handling strategy for the misfire detection area based on the misfire detection results, and to feed back the handling strategy to the staff, so as to realize intelligent identification and handling of blasting misfires based on multimodal perception.

[0015] This invention provides a method and device for intelligent identification and handling of misfires in blasting based on multimodal perception. By collecting multimodal perception data and analyzing the data using deep learning, the detection result of the misfire and the corresponding handling strategy can be determined. This effectively overcomes the limitations of a single sensor in complex blasting sites, significantly reduces the missed detection rate and false detection rate, and eliminates the need for workers to directly enter blasting areas full of unexploded ordnance, greatly reducing the risk of injury or death to workers. This aligns with the development trend of mechanization and automation in mine safety. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart of a method for intelligent identification and handling of misfires based on multimodal perception, proposed in an embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of an intelligent identification and disposal device for misfired explosives based on multimodal perception, as proposed in an embodiment of the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] like Figure 1 As shown, this embodiment of the invention provides a method for intelligent identification and handling of misfires based on multimodal perception, including: S101. After the blasting operation is completed and the prescribed smoke extraction waiting time has elapsed, collect multimodal sensing data corresponding to the blasting operation area; wherein, the multimodal sensing data includes visible light image sequences, infrared thermal imaging image sequences, and lidar point cloud data; For example, after the blasting operation is completed and a prescribed smoke extraction waiting time has elapsed (e.g., 15-30 minutes, to ensure that the concentration of dust and harmful gases on site drops below safety standards), a drone equipped with multimodal sensors enters the blasting area. The sensors include: a high-definition visible light camera to capture high-resolution visible light image sequences of the blasting area; an infrared thermal imager to acquire infrared thermal image sequences of the blasting area and capture temperature distribution; and a lidar system to scan the blasting area and acquire high-precision lidar point cloud data, reflecting the three-dimensional geometric features of the blast pile and boreholes.

[0020] This invention enables non-contact detection after blasting by using drones or robots equipped with sensors for remote data acquisition. Workers no longer need to directly enter the blasting area, which is fraught with the risk of unexploded or misfires, significantly reducing the risk of injury and death for workers. This aligns with the trend of mechanization and automation in mine safety development.

[0021] S102. Use a first deep learning model to extract the first feature corresponding to the visible light image sequence, use a second deep learning model to extract the second feature corresponding to the infrared thermal imaging image sequence, and use a third deep learning model to extract the third feature corresponding to the lidar point cloud data. For example, a first deep learning model can be used to extract the first feature corresponding to the visible light image sequence. In this embodiment, YOLOv5 is chosen as the first deep learning model because it has a good balance between target detection speed and accuracy. Each frame of the visible light image sequence is input into the YOLOv5 model one by one, and the feature map output by the convolutional layer is extracted as the first feature. A second deep learning model is used to extract the second feature corresponding to the infrared thermal imaging image sequence. The YOLOv5 model is also used, but it is fine-tuned using an infrared image dataset during the pre-training stage. Infrared images are input into the model one by one to obtain the second feature. A third deep learning model is used to extract the third feature corresponding to the lidar point cloud data. The PointNet++ model is chosen because it excels at processing point cloud data and can effectively extract local geometric features. The lidar point cloud data is directly input into the PointNet++ model to obtain the third feature.

[0022] S103. The first feature, the second feature and the third feature are fused to obtain a fused feature, and a fourth deep learning model is used to identify the fused feature to obtain the blind shot detection result. For example, the first, second, and third features can be fused at the feature level. Specifically, this can be done using a concatenation operation to connect feature vectors from different modalities, or by using an attention mechanism to weight important features. Alternatively, the first, second, and third features can be concatenated into a single vector.

[0023] The fourth deep learning model can be a fully connected neural network or a support vector machine, or another deep classification network, used to determine whether the fused features belong to a misfire, and the specific type of misfire (such as a misfire inside the borehole or a misfire on the ground). The output of the misfire detection result includes the location coordinates of the misfire (which can be determined based on the coordinates of the multimodal perception data collected by the UAV), the confidence score (i.e., the maximum confidence score in the classification result, the label corresponding to which the maximum confidence score represents the probability of the misfire existing), etc. Therefore, the fourth deep learning model can be trained based on the real misfire detection labels. Training can use the algorithm provided in the embodiments of this application, or it can use existing algorithms.

[0024] For example, the result of a blind shot detection can be: Level 1 Risk (High Risk): Unexploded ordnance has been clearly detected, and the surrounding rock mass is unstable.

[0025] Level 2 Risk (General): Suspected unexploded detonating cord or shallow misfire detected.

[0026] Level 3 Risk (Safety): No abnormalities were found.

[0027] For Level 1 and Level 2 risks, response strategies are automatically generated.

[0028] S104. Based on the results of the misfire detection, formulate a corresponding handling strategy for the misfire detection area and feed the handling strategy back to the staff to realize intelligent identification and handling of blasting misfires based on multimodal perception.

[0029] For Level 1 risks, mechanical detonation can be used. Specifically, for misfired cannons that must be destroyed, a robotic arm is controlled to drop a small shaped charge, and the cannon is detonated remotely after personnel have evacuated to a safe area.

[0030] For secondary risks, water injection / grouting failure can be addressed by precisely aiming the water injection gun at the end at the borehole of the blind blasting gun and injecting high-pressure water or chemical flame retardants to make the explosive lose its explosive sensitivity.

[0031] After the treatment is completed, the control sensors will scan the area again to confirm until the model outputs a safe risk level.

[0032] This invention employs a fusion of three modal data sources: visible light, infrared thermal imaging, and lidar. Visible light images provide rich texture and color information, reflecting the surface condition of the blast hole; infrared thermal imaging can capture abnormally high-temperature areas generated by chemical reactions or friction in unexploded ordnance, compensating for the limitations of visible light at night or in dusty conditions; lidar point cloud data provides precise depth and spatial geometry information, enabling the identification of the 3D morphology of the blast hole and distinguishing between blast pile undulations and the blast hole structure. Through the complementary fusion of multimodal features, this invention effectively overcomes the limitations of single sensors in complex blasting scenarios (such as insufficient lighting, dense smoke, and rock obstruction), significantly reducing the false negative and false positive rates.

[0033] This invention not only identifies misfires but also intelligently matches appropriate response strategies (such as grouting failure or mechanical detonation) based on the detection results (e.g., location, depth, and type of the misfire) and provides feedback to the personnel. This provides scientific decision support for on-site handling, avoids the arbitrariness and danger of relying on experience, shortens accident handling time, and improves the continuity of mine production.

[0034] In this embodiment of the invention, the first deep learning model and the second deep learning model are both set to YOLOv5, CNN (convolutional neural network) or Faster R-CNN (faster region convolutional neural network), and the third deep learning model is set to PointNet++.

[0035] When using deep learning models for blind shot detection, the model's performance is highly dependent on the optimization of its parameters. Traditional gradient descent algorithms are prone to getting stuck in local optima, resulting in poor model generalization ability and insufficient extraction of blind shot features in complex scenarios. Therefore, designing an efficient and robust model parameter optimization algorithm to improve the recognition accuracy of multimodal fusion models is a pressing technical problem. This invention provides a multi-strategy optimization algorithm for training to improve the training effect of the detection model, thereby enhancing the accuracy of blind shot detection.

[0036] In this embodiment of the invention, the first deep learning model, the second deep learning model, and the third deep learning model are all trained using a multi-strategy optimization algorithm before use, including: Use the first deep learning model, the second deep learning model, or the third deep learning model as the target deep learning model; The model parameters of the target deep learning model are initialized and encoded into vectors multiple times to obtain multiple individual model parameters; For example, the model parameters can be randomly initialized between the upper and lower limits and encoded into vectors to obtain individual model parameters.

[0037] Obtain the fitness of each individual model parameter, and determine the optimal individual among the individual model parameters based on the fitness. Fitness can be the reciprocal of the loss function value, which can be obtained through cross-entropy loss or root mean square loss. To avoid the denominator being zero, a very small constant term (such as 0.001) can be added to the denominator. Different deep learning models can be trained using corresponding real data, and the expected labels of different deep learning models can all be real blind shot detection labels, obtained by actual measurements by staff.

[0038] The model parameter individuals are subjected to optimal noise sampling on the optimal individual to obtain the model parameter individuals after optimal noise sampling; Based on the dispersion of individual model parameters in the solution space after optimal noise acquisition, local augmentation decision parameters and global augmentation decision parameters are obtained. Decisions are made based on the local enhancement decision parameters, and the individual model parameters after the optimal noise acquisition are selectively and continuously fitted to obtain the individual model parameters after continuous fitting. Decisions are made based on the global enhancement decision parameters, and the individual model parameters after continuous fitting are selectively mutated to obtain the mutated individual model parameters. Determine whether the number of training iterations has reached the preset maximum number of training iterations. If so, determine the final model parameters of the target deep learning model based on the individual model parameters after the mutation process. Otherwise, based on the target deep learning model, return to the step of obtaining the optimal individual.

[0039] In this embodiment of the invention, the model parameter individuals are subjected to optimal noise acquisition on the optimal individual to obtain the model parameter individuals after optimal noise acquisition, including: The model parameters are used to collect noise from the optimal individual, and the disturbance information is as follows: ; ; In the formula, For the first i The first model parameter of the individual d The perturbation of the dimensional parameter, i =1,2,…,M, where M is the total number of individual model parameters. d =1,2,…,NP, where NP is the total number of parameters in each individual model parameter. Let be the first constant between (0.1, 3). The first random number between (0,1) For intermediate parameters, For the first i The Euclidean distance between each model parameter individual and the optimal individual Pi The second random number between (0,1) The second constant is set to 1 or 2, t is the number of training iterations, and T is the maximum number of training iterations; Based on the disturbance information and the optimal individual, the optimal model parameter individual after noise acquisition is obtained as follows: ; in, The first optimal individual d Dimensional parameters, For the first i The model parameters of the individual after the optimal noise acquisition. d Dimensional parameters, The third random number between (0,1) The fourth random number between (0,1) As an information guiding factor, and , For the first i +1 model parameter individual's first d dimensional parameters, and i When it is M, For the individual parameters of the random modeld Dimensional parameters.

[0040] By introducing the cosine function and the dynamic perturbation of the distance between individuals, we not only increase the diversity of the population and avoid premature convergence, but also adaptively adjust the search step size according to the distance between an individual and the optimal individual, thus balancing the ability of global exploration and local development.

[0041] In this embodiment of the invention, based on the dispersion of individual model parameters in the solution space after optimal noise acquisition, local enhancement decision parameters and global enhancement decision parameters are obtained, including: The dispersion of individual model parameters in the solution space after obtaining the optimal noise acquisition is as follows: ; In the formula, For dispersion parameters, This is the scaling parameter, and it is set to 80. For the first k The fitness of individual model parameters after optimal noise acquisition The average fitness of individual model parameters after all optimal noise acquisitions. k =1,2,…,M; Based on the aforementioned dispersion, the local augmentation decision parameters and the global augmentation decision parameters are determined as follows: ; ; In the formula, To locally enhance decision parameters, To enhance decision parameters globally, For the first h The training progress factor for each successive fitting operation. The total number of times continuous fitting is performed. For the first g Training progress factor for each mutation process. The training progress factor is the total number of mutation operations performed. , Fitness is the fitness level prior to performing continuous fitting or mutation treatment. Max is the function that takes the maximum value after performing continuous fitting or mutation processing.

[0042] This invention utilizes the dispersion of fitness to dynamically adjust the probabilities of local enhancement and global mutation. When the population dispersion is large, the algorithm tends to engage in local expansion to achieve rapid convergence; when the population tends to be uniform (getting trapped in a local optimum), the algorithm automatically increases the mutation probability to escape the trap. This mechanism makes the optimization process more intelligent, eliminating the need for frequent manual parameter adjustments.

[0043] In this embodiment of the invention, decision-making is made based on the local enhancement decision parameters, and the individual model parameters after the optimal noise acquisition are selectively and continuously fitted to obtain the individual model parameters after continuous fitting, including: A first decision factor is randomly generated between (0, 1); if the first decision factor is greater than the local augmentation decision parameter, no operation is performed. When the first decision factor is less than the local enhancement decision parameter, the individual model parameters after the optimal noise acquisition are continuously fitted to obtain the individual model parameters after continuous fitting: ; ; ; In the formula, For the first m Individual model parameters after optimal noise acquisition For the first m Individual model parameters after continuous fitting. The first random individual is randomly selected from the model parameters individuals after optimal noise acquisition. This refers to the second random individual selected from the model parameter individuals after optimal noise acquisition. The third random individual is randomly selected from the model parameter individuals after the optimal noise acquisition. It is a natural constant. For the optimal individual, The first fitting coefficient, The second fitting coefficient, for , as well as The corresponding mean individual, i.e., the parameter of each dimension, is , as well as In the same dimension, the mean of the parameters, for The corresponding advantages and disadvantages of the location, for The corresponding advantages and disadvantages of the location, for The corresponding advantages and disadvantages of the location, The basic weighting coefficient is set to 0.01; Location superiority / inferiority = ; For the first j The fitness of a random individual The fitness of the optimal individual.

[0044] This continuous fitting is a highly intelligent, adaptive, and robust local optimization strategy. Through adaptive decision-making, it ensures that a refined search is performed at the appropriate time, enhancing the directionality and comprehensiveness of the search. It also introduces positional advantages and disadvantages to make full use of the information of other individuals, improving the efficiency and comprehensiveness of spatial exploration, thereby enhancing the ability to find the global optimum and improving the training effect.

[0045] In this embodiment of the invention, decisions are made based on the global enhancement decision parameters, and the individual model parameters after continuous fitting are selectively mutated to obtain mutated individual model parameters, including: A second decision factor is randomly generated between (0, 1); if the second decision factor is greater than the global enhanced decision parameter, no operation is performed. When the second decision factor is less than the global augmentation decision parameter, the individual model parameters after continuous fitting are mutated to obtain the mutated individual model parameters as follows: ; ; In the formula, For the first n The first individual of the model parameters after the nth consecutive fitting d Dimensional parameters, For the first n Individual model parameters after mutation treatment As a variable factor, This is the variation range control factor, and it is set to 0.45; The order of mutation, This represents the total mutation order, and is set to 3. For the first d The upper limit of the dimension parameter.

[0046] Optionally, after each processing of individual model parameters, limit-crossing processing can be performed on these individual parameters to ensure parameter validity. Simulated annealing can also be used for mutation processing to ensure the algorithm's training speed.

[0047] This mutation process is a highly intelligent, adaptive, and robust optimization strategy. It balances exploration and exploitation through adaptive decision-making, enhances global search capabilities by utilizing higher-order differences, achieves efficient convergence with dynamic step size, and ensures parameter validity through boundary constraints. It can significantly improve the model's recognition accuracy, generalization ability, and training efficiency, ultimately making the blind gun recognition system based on multimodal perception more reliable and accurate.

[0048] In this embodiment of the invention, a first deep learning model is used to extract a first feature corresponding to the visible light image sequence, a second deep learning model is used to extract a second feature corresponding to the infrared thermal imaging image sequence, and a third deep learning model is used to extract a third feature corresponding to the lidar point cloud data, including: The visible light images in the visible light image sequence are input one by one into the first deep learning model, and the output of the first deep learning model is used to construct the first feature. The infrared thermal imaging images in the infrared thermal imaging image sequence are input one by one into the second deep learning model, and the output of the second deep learning model is used to construct the second feature. The lidar point cloud data is input into the third deep learning model to obtain the third feature corresponding to the lidar point cloud data.

[0049] In this embodiment of the invention, the disposal strategy includes at least: a grouting failure disposal strategy and a mechanical detonation strategy.

[0050] like Figure 2 As shown, the present invention provides an intelligent identification and handling device for misfires based on multimodal perception, comprising: The multimodal data acquisition module 201 is used to acquire multimodal sensing data corresponding to the blasting operation area after the blasting operation is completed and a specified smoke extraction waiting time has elapsed; wherein, the multimodal sensing data includes visible light image sequences, infrared thermal imaging image sequences, and lidar point cloud data; Feature extraction module 202 is used to extract the first feature corresponding to the visible light image sequence using a first deep learning model, extract the second feature corresponding to the infrared thermal imaging image sequence using a second deep learning model, and extract the third feature corresponding to the lidar point cloud data using a third deep learning model. The blind shot detection module 203 is used to fuse the first feature, the second feature and the third feature to obtain the fused feature, and to use a fourth deep learning model to identify the fused feature to obtain the blind shot detection result. The handling suggestion module 204 is used to formulate a handling strategy for the misfire detection area based on the misfire detection results, and to feed back the handling strategy to the staff, so as to realize intelligent identification and handling of blasting misfires based on multimodal perception.

[0051] The present invention provides an intelligent identification and disposal device for blind blasting based on multimodal perception, which can perform the above-mentioned method and technical solution. Its principle and beneficial effects are similar, and will not be repeated here.

[0052] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0053] Embodiments of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, electronic devices, and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0054] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0055] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0056] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.

[0057] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0058] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for intelligent identification and handling of misfires based on multimodal perception, characterized in that, include: After the blasting operation is completed and the prescribed smoke extraction waiting time has elapsed, multimodal sensing data corresponding to the blasting operation area is collected; wherein, the multimodal sensing data includes visible light image sequences, infrared thermal imaging image sequences, and lidar point cloud data; A first deep learning model is used to extract the first feature corresponding to the visible light image sequence, a second deep learning model is used to extract the second feature corresponding to the infrared thermal imaging image sequence, and a third deep learning model is used to extract the third feature corresponding to the lidar point cloud data. The first feature, the second feature, and the third feature are fused to obtain a fused feature, and a fourth deep learning model is used to identify the fused feature to obtain the blind shot detection result; Based on the results of the misfire detection, a corresponding handling strategy is formulated for the misfire detection area, and the handling strategy is fed back to the staff to realize intelligent identification and handling of blasting misfires based on multimodal perception.

2. The intelligent identification and handling method for misfires based on multimodal perception as described in claim 1, characterized in that, The first and second deep learning models are both set to YOLOv5, CNN or Faster R-CNN, and the third deep learning model is set to PointNet++.

3. The intelligent identification and handling method for misfires based on multimodal perception as described in claim 1, characterized in that, Before using the first, second, and third deep learning models, a multi-strategy optimization algorithm was employed for training, including: Use the first deep learning model, the second deep learning model, or the third deep learning model as the target deep learning model; The model parameters of the target deep learning model are initialized and encoded into vectors multiple times to obtain multiple individual model parameters; Obtain the fitness of each individual model parameter, and determine the optimal individual among the individual model parameters based on the fitness. The model parameter individuals are subjected to optimal noise sampling on the optimal individual to obtain the model parameter individuals after optimal noise sampling; Based on the dispersion of individual model parameters in the solution space after optimal noise acquisition, local augmentation decision parameters and global augmentation decision parameters are obtained. Decisions are made based on the local enhancement decision parameters, and the individual model parameters after the optimal noise acquisition are selectively and continuously fitted to obtain the individual model parameters after continuous fitting. Decisions are made based on the global enhancement decision parameters, and the individual model parameters after continuous fitting are selectively mutated to obtain the mutated individual model parameters. Determine whether the number of training iterations has reached the preset maximum number of training iterations. If so, determine the final model parameters of the target deep learning model based on the individual model parameters after the mutation process. Otherwise, based on the target deep learning model, return to the step of obtaining the optimal individual.

4. The intelligent identification and handling method for misfires based on multimodal perception as described in claim 3, characterized in that, The model parameter individuals are subjected to optimal noise sampling on the optimal individual to obtain the model parameter individuals after optimal noise sampling, including: The model parameters are used to collect noise from the optimal individual, and the disturbance information is as follows: ; ; In the formula, For the first i The first model parameter of the individual d The perturbation of the dimensional parameter, i =1,2,…,M, where M is the total number of individual model parameters. d =1,2,…,NP, where NP is the total number of parameters in each individual model parameter. Let be the first constant between (0.1, 3). The first random number between (0,1) For intermediate parameters, For the first i The Euclidean distance between each model parameter individual and the optimal individual Pi The second random number between (0,1) The second constant is set to 1 or 2, t is the number of training iterations, and T is the maximum number of training iterations; Based on the disturbance information and the optimal individual, the optimal model parameter individual after noise acquisition is obtained as follows: ; in, The first optimal individual d Dimensional parameters, For the first i The model parameters of the individual after the optimal noise acquisition. d Dimensional parameters, The third random number between (0,1) The fourth random number between (0,1) As an information guiding factor, and , For the first i +1 model parameter individual's first d dimensional parameters, and i When it is M, For the individual parameters of the random model d Dimensional parameters.

5. The intelligent identification and handling method for misfires based on multimodal perception according to claim 4, characterized in that, Based on the dispersion of individual model parameters in the solution space after optimal noise acquisition, local augmentation decision parameters and global augmentation decision parameters are obtained, including: The dispersion of individual model parameters in the solution space after obtaining the optimal noise acquisition is as follows: ; In the formula, For dispersion parameters, This is the scaling parameter, and it is set to 80. For the first k The fitness of individual model parameters after optimal noise acquisition The average fitness of individual model parameters after all optimal noise acquisitions. k =1,2,…,M; Based on the aforementioned dispersion, the local augmentation decision parameters and the global augmentation decision parameters are determined as follows: ; ; In the formula, To locally enhance decision parameters, To enhance decision parameters globally, For the first h The training progress factor for each successive fitting operation. The total number of times continuous fitting is performed. For the first g Training progress factor for each mutation process. The training progress factor is the total number of mutation operations performed. , The fitness level before performing continuous fitting or mutation treatment, Max is the function that takes the maximum value after performing continuous fitting or mutation processing.

6. The intelligent identification and handling method for misfires based on multimodal perception as described in claim 3, characterized in that, Decisions are made based on the local enhancement decision parameters, and the individual model parameters after the optimal noise acquisition are selectively and continuously fitted to obtain the individual model parameters after continuous fitting, including: The first decision factor is randomly generated between (0, 1); When the first decision factor is less than the local enhancement decision parameter, the individual model parameters after the optimal noise acquisition are continuously fitted to obtain the individual model parameters after continuous fitting: ; ; ; In the formula, For the first m Individual model parameters after optimal noise acquisition For the first m Individual model parameters after continuous fitting. The first random individual is randomly selected from the model parameters individuals after optimal noise acquisition. This refers to the second random individual selected from the model parameter individuals after optimal noise acquisition. The third random individual is randomly selected from the model parameter individuals after the optimal noise acquisition. It is a natural constant. For the optimal individual, The first fitting coefficient, The second fitting coefficient, for , as well as The corresponding mean individual, for The corresponding advantages and disadvantages of the location, for The corresponding advantages and disadvantages of the location, for The corresponding advantages and disadvantages of the location, The basic weighting coefficient is set to 0.01; Location superiority / inferiority = ; For the first j The fitness of a random individual The fitness of the optimal individual.

7. The intelligent identification and handling method for misfires based on multimodal perception according to claim 3, characterized in that, Decisions are made based on the global enhancement decision parameters, and the individual model parameters after continuous fitting are selectively mutated to obtain mutated individual model parameters, including: A second decision factor is randomly generated between (0, 1); When the second decision factor is less than the global augmentation decision parameter, the individual model parameters after continuous fitting are mutated to obtain the mutated individual model parameters as follows: ; ; In the formula, For the first n The first individual of the model parameters after the nth consecutive fitting d Dimensional parameters, For the first n Individual model parameters after mutation treatment As a variable factor, This is the variation range control factor, and it is set to 0.45; The order of mutation, This represents the total mutation order, and is set to 3. For the first d The upper limit of the dimension parameter.

8. The intelligent identification and handling method for misfires based on multimodal perception as described in claim 1, characterized in that, A first feature corresponding to the visible light image sequence is extracted using a first deep learning model; a second feature corresponding to the infrared thermal imaging image sequence is extracted using a second deep learning model; and a third feature corresponding to the lidar point cloud data is extracted using a third deep learning model, including: The visible light images in the visible light image sequence are input one by one into the first deep learning model, and the output of the first deep learning model is used to construct the first feature. The infrared thermal imaging images in the infrared thermal imaging image sequence are input one by one into the second deep learning model, and the output of the second deep learning model is used to construct the second feature. The lidar point cloud data is input into the third deep learning model to obtain the third feature corresponding to the lidar point cloud data.

9. The intelligent identification and handling method for misfires based on multimodal perception as described in claim 1, characterized in that, The disposal strategies include at least: grouting failure disposal strategies and mechanical detonation strategies.

10. A multimodal perception-based intelligent identification and handling device for misfires, wherein the multimodal perception-based intelligent identification and handling device for misfires is used to execute the multimodal perception-based intelligent identification and handling method for misfires as described in any one of claims 1 to 9, characterized in that, include: The multimodal data acquisition module is used to acquire multimodal sensing data corresponding to the blasting operation area after the blasting operation is completed and a specified smoke extraction waiting time has elapsed; wherein, the multimodal sensing data includes visible light image sequences, infrared thermal imaging image sequences, and lidar point cloud data; The feature extraction module is used to extract the first feature corresponding to the visible light image sequence using a first deep learning model, extract the second feature corresponding to the infrared thermal imaging image sequence using a second deep learning model, and extract the third feature corresponding to the lidar point cloud data using a third deep learning model. The blind shot detection module is used to fuse the first feature, the second feature and the third feature to obtain the fused feature, and to use a fourth deep learning model to identify the fused feature to obtain the blind shot detection result. The handling suggestion module is used to formulate a handling strategy for the misfire detection area based on the misfire detection results, and to feed back the handling strategy to the staff, so as to realize intelligent identification and handling of blasting misfires based on multimodal perception.