Phosphorite underground hidden danger identification method and system based on machine vision

By acquiring multimodal data of phosphate mines through intelligent video sensors, performing spatiotemporal alignment and attention linear processing, and combining dilated convolution and residual networks to generate image features with consistent brightness, and using a multi-sub-discriminator for comprehensive processing, the problem of low accuracy in identifying hidden dangers in phosphate mines has been solved, achieving accurate identification of hidden dangers and safety assurance.

CN121505504APending Publication Date: 2026-02-10GUIZHOU FULIN MINING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511467257.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-14
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

In the underground transportation process of phosphate mines, low light intensity, high dust levels, and fog interference affect the accuracy of identification, resulting in low precision in the process of identifying potential hazards.

Method used

Phosphate mine data is acquired through intelligent video sensors, and spatiotemporal alignment of multimodal data is performed. Attention linear processing and dilated convolution with residual networks are used to generate image features with consistent brightness. Finally, multiple sub-discriminators are used for comprehensive processing to generate hazard identification results.

Benefits of technology

It enables comprehensive and accurate identification of hidden dangers in phosphate mines under complex environments, ensuring production safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121505504A_ABST
    Figure CN121505504A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of electric data processing, and provides a phosphorite underground hidden danger identification method and system based on machine vision. Phosphorite data are obtained through an intelligent video sensor; carrying out alignment processing on the multi-modal phosphorite data to generate a time-space aligned data matrix; performing attention-based linear processing on the spatial dimension corresponding to the data matrix to obtain a feature matrix; processing the feature matrix through preset hole convolution and a residual network to generate image features with consistent brightness; and inputting the image features into a preset number of sub discriminators, respectively outputting standby discrimination results, and comprehensively processing the standby discrimination results to generate a hidden danger discrimination result. Key features are focused through linear processing based on attention, brightness consistency is guaranteed through cavity convolution and a residual network, and feature quality is improved. And finally, the sub-discriminator comprehensively outputs a hidden danger result, and the hidden danger in the phosphorite well is comprehensively and accurately recognized through multi-link cooperation, so that the production safety is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of electric data processing, in particular to a phosphor mine underground hidden danger identification method and system based on machine vision. BACKGROUND

[0002] It is very easy to collect such clear paired data sets in normal environments, and paired data sets have significant advantages in image enhancement tasks, especially when image restoration is required. Paired data sets provide clear supervision signals and accurate learning goals. Paired data sets provide clear inputs for the model, and the output corresponds to the relationship. In the low-light image enhancement task, the paired data set contains images under low-light conditions and corresponding high-light images. The model can learn how to restore details, brightness and contrast in the image by comparing the differences between the input and the target image. Paired data sets can also help accurately calculate loss functions, especially pixel-level loss calculations.

[0003] In the scene of foreign matter detection in the transportation link of the phosphor mine, low illumination, high dust, and fog interference are the core problems affecting the recognition accuracy. The standard generator structure is easy to lose key details such as texture and edge in the encoding process, resulting in the inability of the enhanced image to clearly present small foreign objects (such as bolts and tire fragments) or fine cracks. The model has limited brightness enhancement capability for extremely dark areas in the mine, and lacks pertinence, making it difficult to balance overall brightness and local overexposure. Therefore, the above situations are prone to cause the problem of low accuracy in the process of identifying hidden dangers in the phosphor mine. SUMMARY

[0004] The present application provides a phosphor mine underground hidden danger identification method and system based on machine vision, which can at least partially solve the problem of low accuracy in the process of identifying hidden dangers in the phosphor mine.

[0005] Other characteristics and advantages of the present application will become apparent from the following detailed description, or will be learned by practice of the present application.

[0006] According to one aspect of the present application, a phosphor mine underground hidden danger identification method based on machine vision is provided, comprising: acquiring phosphor mine data through an intelligent video sensor; aligning the multi-modal phosphor mine data to generate a spatio-temporal aligned data matrix; performing attention-based linear processing on the spatial dimension corresponding to the data matrix to obtain a feature matrix; processing the feature matrix through a pre-set hollow convolution and residual network to generate an image feature with consistent brightness; inputting the image feature into a pre-set number of sub-discriminators to respectively output standby discrimination results, and performing comprehensive processing on the standby discrimination results to generate a hidden danger discrimination result.

[0007] In the present application, based on the foregoing scheme, the alignment processing of the multi-modal phosphate rock data generates a spatio-temporal aligned data matrix, including: filtering processing of the multi-modal phosphate rock data based on the energy gradient of the phosphate rock data to generate first data; multi-modal data alignment based on dynamic time warping is performed on the first data to generate a spatio-temporal aligned data matrix.

[0008] In the present application, based on the foregoing scheme, the alignment processing of the multi-modal phosphate rock data generates a spatio-temporal aligned data matrix, including: filtering processing of the multi-modal phosphate rock data based on the energy gradient of the phosphate rock data to generate first data; multi-modal data alignment based on dynamic time warping is performed on the first data to generate a spatio-temporal aligned data matrix.

[0009] In the present application, based on the foregoing scheme, the alignment processing of the multi-modal phosphate rock data generates a spatio-temporal aligned data matrix, including: filtering processing of the multi-modal phosphate rock data based on the energy gradient of the phosphate rock data to generate first data; multi-modal data alignment based on dynamic time warping is performed on the first data to generate a spatio-temporal aligned data matrix. c wherein, H, W, C respectively represent height, width and channel number, i, j, c respectively represent height identifier, width identifier and channel identifier.

[0010] In the present application, based on the foregoing scheme, the alignment processing of the multi-modal phosphate rock data generates a spatio-temporal aligned data matrix, including: filtering processing of the multi-modal phosphate rock data based on the energy gradient of the phosphate rock data to generate first data; multi-modal data alignment based on dynamic time warping is performed on the first data to generate a spatio-temporal aligned data matrix. c wherein, represents convolution kernel weight, represents channel c statistical information, c and C represent channel identification and total number of channels, respectively H, W, C ; respectively represent mapping processing and activation processing.

[0011] ​​​​In this application, based on the aforementioned scheme, the step of processing the feature matrix through a preset dilated convolution and residual network to generate image features with consistent brightness includes: performing dilated convolution and activation processing on the feature matrix to generate activation features; and performing residual connections on the activation features based on the residual network to generate image features with consistent brightness.

[0012] In this application, based on the aforementioned scheme, the step of inputting the image features into a preset number of sub-discriminators, outputting backup discrimination results respectively, and comprehensively processing the backup discrimination results to generate a hazard discrimination result includes: inputting the image features into a preset number of sub-discriminators, outputting backup discrimination results respectively; and comprehensively processing the backup discrimination results based on the discrimination parameters preset for the sub-discriminators to generate a hazard discrimination result.

[0013] According to one aspect of this application, a machine vision-based underground hazard identification system for phosphate mines is provided, comprising: The acquisition module acquires phosphate rock data through intelligent video sensors; The alignment module is used to align the multimodal phosphate rock data to generate a spatiotemporally aligned data matrix. The linear module is used to perform attention-based linear processing on the spatial dimensions corresponding to the data matrix to obtain the feature matrix; The extraction module is used to process the feature matrix through a preset dilated convolution and residual network to generate image features with consistent brightness. The judgment module is used to input the image features into a preset number of sub-discriminators, output backup judgment results respectively, and perform comprehensive processing on the backup judgment results to generate hidden danger judgment results.

[0014] According to one aspect of this application, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processor, implements the machine vision-based method for identifying hidden dangers in phosphate mines as described in the above embodiments.

[0015] According to one aspect of this application, an electronic device is provided, comprising: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the machine vision-based method for identifying hidden dangers in phosphate mines as described in the above embodiments.

[0016] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the machine vision-based method for identifying hidden dangers in underground phosphate mines provided in the various optional implementations described above.

[0017] In the technical solution of this application, phosphate mine data is acquired through an intelligent video sensor; the multimodal phosphate mine data is aligned to generate a spatiotemporally aligned data matrix; attention-based linear processing is performed on the spatial dimension corresponding to the data matrix to obtain a feature matrix; the feature matrix is ​​processed through a preset dilated convolution and residual network to generate image features with consistent brightness; the image features are input into a preset number of sub-discriminators, each outputting backup discrimination results; the backup discrimination results are comprehensively processed to generate a hazard identification result. The acquisition of multimodal phosphate mine data through an intelligent video sensor, followed by alignment processing to ensure precise spatiotemporal correspondence, lays the foundation for accurate analysis. Attention-based linear processing focuses on key features, while dilated convolution and residual networks ensure consistent brightness and improve feature quality. Finally, the sub-discriminators comprehensively output hazard results, enabling multi-stage collaboration for comprehensive and accurate identification of underground phosphate mine hazards, ensuring production safety.

[0018] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0020] i, j, c The flowchart illustrating a machine vision-based method for identifying hidden dangers in underground phosphate mines is shown in one embodiment of this application.

[0021] c and C represent channel identification and total number of channels, respectively The flowchart illustrating the generation of the feature matrix is ​​shown in one embodiment of this application.

[0022] Figure 1 The illustration shows a schematic diagram of a machine vision-based underground hazard identification system for phosphate mines in one embodiment of this application.

[0023] Figure 2A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown. Detailed Implementation

[0024] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.

[0025] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.

[0026] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0027] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0028] The implementation details of the technical solution of this application are described below: Figure 3 A flowchart of a machine vision-based method for identifying hidden dangers in underground phosphate mines according to an embodiment of this application is shown. (Refer to...) Figure 4 As shown, the machine vision-based method for identifying hidden dangers in underground phosphate mines includes at least steps S110 to S150, which are described in detail below: The S110 acquires phosphate rock data through a smart video sensor.

[0029] In one embodiment of this application, before deploying intelligent video sensors in a phosphate mine, targeted configuration based on environmental characteristics is required. First, the sensor's optical parameters, such as focal length, aperture, and exposure time, are adjusted to ensure clear image capture even in low-light or dusty environments. Second, the sensor's sampling frequency is set, balancing real-time performance with data volume; for example, acquiring 25-30 frames per second to capture details of dynamic hazards (such as equipment loosening or rockfall). Finally, the data transmission protocol is configured, selecting a wireless or wired method with strong anti-interference capabilities. A self-organizing network or wireless sensor network is established between the intelligent video sensors to ensure stable data transmission to subsequent processing units.

[0030] Intelligent video sensors integrate multiple sensing modules and sensor chips, including infrared thermal imaging, laser ranging, or acoustic sensors in addition to visible light imaging. During execution, all modules are triggered synchronously to ensure that multimodal data is aligned in time. For example, when a visible light camera captures a frame, the infrared module simultaneously records the temperature distribution of the corresponding area, and the laser module measures spatial distance. This synchronous acquisition avoids data misalignment caused by time differences, providing a foundation for subsequent multimodal fusion.

[0031] In practical applications, the complex environment of phosphate mines, including uneven lighting and dust obstruction, can degrade data quality. The sensor can analyze scene characteristics in real time and dynamically adjust the acquisition strategy. For example, if a localized area is detected as too dark, the algorithm will increase the exposure compensation for that area; if dust is detected causing image blur, dehazing processing will be initiated, estimating and eliminating dust interference by comparing the sharpness differences of adjacent frames. Furthermore, to address image jitter caused by vibration or impact, the sensor will activate electronic image stabilization, outputting a smooth video stream through inter-frame motion estimation and compensation.

[0032] Optionally, the collected multimodal data volume is enormous, and direct transmission would consume a significant amount of bandwidth. The edge computing unit built into the sensor performs preliminary screening, eliminating invalid or duplicate data. For example, areas that do not show significant changes in consecutive frames (such as a stationary rock wall) are marked as low priority, and only keyframes are retained; for temperature data, if the temperature in a certain area remains stable within a safe range for a long period, its sampling frequency is reduced. The filtered data is then compressed using lightweight algorithms (such as JPEG2000 or H.264) to reduce its size while retaining hazard-related features (such as crack edges and abnormal temperature rise points), ensuring the efficiency of industrial big data acquisition and processing.

[0033] In the process described above, the intelligent video sensor possesses multiple sensing capabilities, simultaneously acquiring multimodal data such as visible light images and infrared thermal images from underground phosphate mines. This data reflects environmental information from different perspectives; for example, visible light images clearly show the appearance and shape of objects, while infrared thermal images reflect the temperature distribution of objects. The complementary nature of these multiple modal data provides rich raw information for subsequent comprehensive and accurate identification of potential hazards.

[0034] S120, The multimodal phosphate rock data is aligned to generate a spatiotemporally aligned data matrix.

[0035] In practical applications, multimodal phosphate rock data is acquired synchronously or asynchronously by different sensors, such as visible light cameras, infrared thermal imagers, and lidar. Differences in sampling frequency and startup time among these sensors can lead to timeline misalignment. The first step in alignment is to establish a unified time reference. A reference sensor, such as a camera with a high-precision clock synchronization, is selected as the time anchor. By recording the absolute timestamps of data acquisition from other sensors, all modal data are mapped onto the same timeline. For example, if the infrared sensor starts up 0.5 seconds later than the camera, the timestamp of the infrared data is shifted forward by 0.5 seconds to ensure that the multimodal data is aligned in the time dimension, avoiding spatial feature mismatch due to time differences.

[0036] Different sensors may have different spatial reference systems. For example, cameras use image pixels as coordinates, while LiDAR uses 3D point clouds. Spatial registration transforms all data to the same coordinate system. First, the spatial position and orientation of each sensor are calibrated, such as using a checkerboard pattern to calibrate the camera's intrinsic parameters. The data is then used to correct the LiDAR's pose. Finally, geometric transformations, such as rotation, translation, and scaling, are used to map the data to a unified 3D space or 2D image plane. For instance, the coordinates of an obstacle in the LiDAR point cloud are projected onto the corresponding pixel position in the camera image through perspective transformation, ensuring that the multimodal data points to the same physical region in spatial dimensions.

[0037] Optionally, even after time standardization, local temporal misalignments may still exist between different modal data due to differences in sampling frequencies, such as 30 frames per second for a camera and 10 frames per second for infrared. A dynamic time warping algorithm is employed to find the optimal local alignment path by constructing a distance matrix between time series. For example, by comparing the data features of three consecutive camera frames with one infrared frame, such as edge intensity and average temperature, if the second frame of camera data shows the highest feature similarity to the infrared frame, then the infrared frame is aligned to the second frame position on the camera timeline. This process compensates for detail misalignments caused by inconsistent sampling frequencies by dynamically adjusting the temporal correspondence.

[0038] In one embodiment of this application, the multimodal phosphate rock data is aligned to generate a spatiotemporally aligned data matrix, including: Based on the energy gradient of the phosphate rock data, the multimodal phosphate rock data is filtered to generate the first data. The first data is subjected to multimodal data alignment based on dynamic time warping to generate a spatiotemporally aligned data matrix.

[0039] In one embodiment of this application, visible light images are treated with wavelet threshold shrinkage to remove dust interference. An adaptive filter based on wavelet transform and attention mechanism is designed to address image degradation issues in high-dust, low-light environments underground. Energy weights are determined based on the energy gradient of phosphate rock data. for: in, Figure 1 The x and y axes represent the phosphate rock data. N This indicates the total number of phosphate rock data. Represents the energy gradient, used to dynamically adjust the filter strength; E and F These represent the energy and numerical values ​​of phosphate rock data, respectively.

[0040] In one embodiment of this application, after determining the energy weight, the multimodal phosphate rock data is analyzed based on the energy weight. Filter the data to generate the first data. for: in, W This indicates the preset threshold shrinkage operator. Indicates in The data value of the phosphate rock at the point.

[0041] The above process reflects the degree of change in the data through the energy gradient. By filtering based on the energy gradient, noise and interference information in the data can be removed, while retaining key features related to potential hazards. For example, in infrared thermal images, filtering can remove abnormal temperature fluctuations caused by sensor noise or environmental factors, making the true temperature change characteristics more prominent and improving the quality and reliability of the data.

[0042] In one embodiment of this application, a multimodal data alignment algorithm based on dynamic time warping is proposed to align the acquired multimodal data. The first data is then aligned using dynamic time warping to generate a spatiotemporally aligned data matrix. D for: in, Representing modes i and j Inter-correlation weights, Indicates the first i Visual modal features, such as camera pixel values. Indicates the first j indivual t Non-visual modal features at any given time, such as LiDAR point cloud coordinates, Indicates dynamic weights. The total variational regularization term representing time warp, Indicates the time factor. N and M These represent the total number of modes.

[0043] Through the above preprocessing steps, a spatiotemporally aligned multimodal data cube is output. A dynamic scene model of the downhole environment is constructed to address the problem that traditional single-vision sensors are susceptible to interference from dust and changes in lighting.

[0044] After spatiotemporal alignment, the multimodal data is fused according to the aligned time and spatial coordinates to generate a structured data matrix. The matrix's dimensional design balances computational efficiency and feature completeness, typically employing a three-dimensional structure of time, space, and modality. For example, each row of the matrix represents a time point, each column represents a spatial location, and different channels store data for different modalities, such as visible light intensity, infrared temperature, and laser distance. During fusion, the validity of the data is checked. If data for a certain modality is missing at a certain time-space point, it is filled in through neighborhood interpolation or intermodal correlation prediction. The final generated data matrix reflects the spatiotemporal dynamics of the phosphate mine scene while preserving the complementarity of multimodal features, providing a unified input for subsequent processing.

[0045] The above process, through a dynamic time warping algorithm, automatically adjusts the time series of different modalities, aligning them in the time dimension. Simultaneously, by incorporating spatial location information, the multimodal data is mapped to a unified spatial coordinate system. This generated spatiotemporally aligned data matrix ensures accurate correspondence between different modalities, providing a reliable foundation for subsequent comprehensive analysis and avoiding misjudgments caused by spatiotemporal inconsistencies.

[0046] S130, perform attention-based linear processing on the spatial dimension corresponding to the data matrix to obtain the feature matrix.

[0047] In one embodiment of this application, a global aggregation operation is first performed on the spatial dimensions (height and width) of the data matrix. Average pooling is then used to compress all pixel values ​​within the spatial range of each channel into a single scalar statistical value. This process is similar to compressing the image for each channel to extract its overall feature intensity. For example, if a channel corresponds to a temperature modality, average pooling calculates the average temperature across all spatial locations of that channel, reflecting the average temperature level of the entire region. This global statistical approach eliminates interference from spatial details, highlights the contribution of each modality to the overall scene, and provides a basis for subsequent attention allocation.

[0048] like Figure 1 As shown, in one embodiment of this application, attention-based linear processing is performed on the spatial dimension corresponding to the data matrix to obtain a feature matrix, including: S210, Perform average pooling on the data matrix in the spatial dimension to obtain the statistical information of the channels; S220, perform convolution and activation processing on the statistical information to generate channel parameters; S230, perform linear processing on the channel dimensions based on the channel parameters to obtain an attention-based feature matrix.

[0049] In one embodiment of this application, to reduce the loss of details in the generator network feature extraction, a channel attention mechanism is introduced into the generator network to improve the model's feature extraction capability. This maintains a significant performance improvement while reducing computational complexity. The data matrix is ​​then subjected to average pooling in the spatial dimension to obtain the channel... c Statistical information for: in, u, v These represent the height, width, and number of channels, respectively. Figure 2 These represent the height, width, and channel identifiers, respectively. Average pooling compresses data in the spatial dimension, calculating statistical information such as the average value for each channel within the spatial range. This process removes interference from spatial details and extracts the global features of each channel. For example, when processing visible light image channels, average pooling can obtain statistical features such as the overall brightness and contrast of that channel, providing a basis for subsequent determination of channel importance.

[0050] Based on global statistics, the correlation between channels is modeled through convolutional operations. The convolutional kernel learns the dependencies between statistical values ​​of different channels; for example, the statistical values ​​of high-temperature regions may be correlated with the statistical values ​​of vibration modes. After convolution, a nonlinear activation function (such as Sigmoid or ReLU) is applied to generate channel parameters, which represent the contribution weight of each channel to the final feature. For example, if the parameter value of a certain channel is close to 1, it means that the feature of that channel is more important for identifying potential hazards in phosphate mines (such as cracks, abnormal temperature rises); if the parameter value is close to 0, it means that the feature of that channel can be suppressed. This process achieves adaptive feature selection by learning the dynamic relationships between channels.

[0051] In one embodiment of this application, the statistical information is convolved and then activated to generate channels. c Channel parameters for: in, Indicates the convolution kernel weights. Indicates channel c Statistical information c and C These represent the channel identifier and the total number of channels, respectively. These represent mapping and activation processing, respectively. Convolution processing can further uncover patterns and relationships in statistical information by learning the correlations between channels through different convolution kernels. Activation processing introduces non-linear factors, enabling the generated channel parameters to more accurately reflect the contribution of channels to the final features. For example, some channels may be closely related to specific types of hazards in phosphate mines; after convolution and activation processing, the parameters of these channels will be assigned higher values, highlighting their importance in hazard identification.

[0052] Optionally, the generated channel parameters are mapped to attention weights, and these weights are multiplied by the corresponding channels of the original data matrix to complete feature recalibration. For example, if a channel has a high weight, the pixel values ​​at all its spatial locations will be amplified, highlighting the features of that channel; conversely, the features of channels with low weights will be suppressed. This attention-based linear processing can dynamically adjust the contribution of each modality, making the model focus more on features related to potential hazards (such as high-temperature anomalies and structural deformation areas) while reducing interference from irrelevant information (such as a uniform background). The recalibrated data matrix retains the original structure in the spatial dimension, but the feature intensity has been redistributed according to the channel importance.

[0053] In one embodiment of this application, the channel dimensions are linearly processed based on the channel parameters to obtain an attention-based feature matrix. The calculated attention weights are then applied to the channel dimensions of the input feature map through a multiplication operation to obtain a weighted feature map.

[0054] After attention recalibration, the processed data matrix is ​​output as the feature matrix. At this point, the feature matrix not only retains the spatial details of the multimodal data but also enhances the feature representation of key modalities through an attention mechanism. For example, in a phosphate mine scenario, the feature matrix may simultaneously contain structural information from visible light images, temperature anomalies from infrared thermal images, and distance changes from lidar, but temperature and structure-related channels are assigned higher weights. This fusion and optimization of multimodal features allows subsequent processing (such as hazard classification or location) to more efficiently utilize complementary information, improving recognition accuracy and robustness. The entire process, by simulating the attention mechanism of human vision, achieves intelligent perception of complex phosphate mine scenarios.

[0055] The above process involves linearly weighting the channel dimensions based on the generated channel parameters. Channel features with high parameters are enhanced, while those with low parameters are suppressed. This attention-based approach allows the model to focus more on key channel features relevant to potential hazards, ignoring irrelevant information and thus improving the feature matrix's ability to represent hazards. In image enhancement tasks in phosphate mines, this method can effectively enhance important features while suppressing irrelevant features, thereby improving image detail and quality, especially in environments like phosphate mines with insufficient lighting, noise interference, and blurred details.

[0056] S140, the feature matrix is ​​processed by a preset dilated convolution and residual network to generate image features with consistent brightness.

[0057] First, the feature matrix is ​​processed hierarchically using pre-defined dilated convolutions. The core of dilated convolution lies in the introduction of spacing into its kernels, which allows the kernels to cover a larger receptive field as they slide, without increasing the number of parameters. For example, the first layer of dilated convolution might use a small dilation rate to capture the fine structure of local regions in the feature matrix, such as tiny cracks or texture variations on the surface of phosphate rock. Subsequent layers gradually increase the dilation rate to expand the receptive field and extract more global features, such as regional temperature anomalies or structural deformation patterns. This multi-scale perception mechanism, through combinations of convolution kernels with different dilation rates, achieves hierarchical feature extraction from local to global, ensuring that both details and overall trends related to brightness changes are effectively captured.

[0058] In one embodiment of this application, the feature matrix is ​​processed using a pre-defined dilated convolution and residual network to generate image features with consistent brightness, including: The feature matrix is ​​subjected to dilated convolution and activation processing to generate activation features; The activation features are residually connected using a residual network to generate image features with consistent brightness.

[0059] In practical applications, due to the poor image quality, insufficient brightness, and blurred details in phosphate mines, the original generator tends to ignore details when processing such images, resulting in blurry images. Therefore, a brightness enhancement module is introduced into the generator, which mainly consists of dilated convolution and residual networks.

[0060] In this embodiment, dilated convolution can use a 3×3 dilated convolution kernel, and different dilation rates are set to expand the receptive field, thereby obtaining more contextual information without increasing computational cost. The purpose of dilated convolution is to preserve image details while enhancing brightness features. Through multiple layers of dilated convolution, the brightness enhancement module can extract and amplify brightness features layer by layer, forming a progressive brightness enhancement effect. Even in the dark areas of a phosphate mine, the module can gradually increase brightness, making the generated image more uniform and clearer. Batch normalization and ReLU activation are performed immediately after each layer of dilated convolution to accelerate training and improve the model's nonlinear representation capability. The features after dilated convolution and activation are added to the input features to achieve residual connections of features, which can avoid information loss and maintain brightness consistency across different resolutions.

[0061] Optionally, a dynamic weight adjustment module is embedded in the residual network to adaptively enhance brightness-related features based on the content of the feature matrix. This module analyzes the feature distribution of the output of each dilated convolution layer, identifying channels or spatial regions closely related to brightness changes. For example, if the feature values ​​of a certain channel show significant differences in the high-temperature region of phosphate rock, the dynamic weight module will assign it a higher weight, making it dominant in subsequent processing; conversely, channels that are not sensitive to brightness changes will be suppressed. This adaptive adjustment mechanism ensures that the network can focus on key regions with inconsistent brightness, such as abnormal temperature rise points or boundaries of sudden illumination changes, while filtering out irrelevant background noise.

[0062] After multi-level dilated convolution and residual processing, the final feature matrix undergoes brightness normalization. Normalization aims to eliminate absolute brightness differences between different modalities or regions, mapping the feature matrix to a uniform brightness range. For example, the global brightness mean and standard deviation of the feature matrix are calculated, and then all pixel values ​​are adjusted to a preset range, such as 0 to 1, through a linear transformation. This process ensures consistency in the brightness dimension of the generated image features, avoiding feature imbalance caused by excessive brightness differences in the original data. The final output feature matrix retains key information about brightness variations in the phosphate mine scene, such as the shadow contrast of cracks and the thermal radiation differences due to temperature anomalies, while achieving stable representation across scenes and modalities through normalization.

[0063] S150, the image features are input into a preset number of sub-discriminators, and backup discrimination results are output respectively. The backup discrimination results are comprehensively processed to generate hidden danger discrimination results.

[0064] In practical applications, the original discriminator typically focuses only on the overall image quality, neglecting local details. Because images from underground phosphate mines often contain significant noise, uneven local lighting, and blurred details, a single discriminator may not be able to adequately capture this detailed information, leading to distortion or loss of detail in certain local areas of the generated image. Furthermore, the original discriminator often cannot effectively capture detail differences at different scales. In underground phosphate mines, images may contain information at multiple scales, from the global scene to small objects; the original discriminator cannot adapt to features at different scales in such cases, causing the generator to fail to accurately recover details.

[0065] In one embodiment of this application, normalized image features are input into a preset number of sub-discriminators, each of which independently receives the complete image feature matrix. During the initialization phase, each sub-discriminator is configured with different parameter sets according to a preset strategy. For example, some sub-discriminators focus on extracting texture features, such as gradient changes at crack edges; some focus on temperature anomalies, such as the distribution pattern of thermal radiation; and some focus on structural deformation, such as interruptions in spatial continuity. This differentiated initialization allows the sub-discriminators to analyze image features from different perspectives, avoiding the limitations of a single discriminator. For example, in a phosphate mine scenario, the first sub-discriminator might capture minute cracks through high-frequency filtering, the second sub-discriminator might identify large-area collapse risks through low-frequency analysis, and the third sub-discriminator might determine comprehensive hidden dangers through multimodal fusion.

[0066] In one embodiment of this application, the image features are input into a preset number of sub-discriminators, which output backup discrimination results respectively. The backup discrimination results are then comprehensively processed to generate a hazard discrimination result, including: The image features are input into a preset number of sub-discriminators, and alternative discrimination results are output respectively; Based on the preset discrimination parameters of the sub-discriminator, the backup discrimination results are comprehensively processed to generate hidden danger discrimination results.

[0067] For example, in this embodiment, a predetermined number of sub-discriminators are selected, such as a dual discriminator. The dual discriminator consists of two structurally identical sub-discriminators, D1 and D2. D1 processes images with a resolution of 256×256, and D2 processes images with a resolution of 128×128. The two sub-discriminators process images of different resolutions to achieve multi-scale feature extraction. This design helps the discriminator capture details in high-resolution images while also taking into account the overall structural features of low-resolution images. Each convolutional layer is followed by an activation function, which can effectively capture local information such as edges and textures in the image. These convolutional layers, with a stride of 2, further reduce the resolution and extract higher-level features.

[0068] For example, in this implementation, each sub-discriminator performs deep analysis on the input image features. The analysis process typically includes multi-level feature extraction and pattern matching: low-level processing (such as convolution operations) focuses on local details (such as brightness abrupt changes in a single pixel), mid-level processing (such as pooling operations) integrates neighborhood information (such as the average temperature of the region), and high-level processing (such as fully connected layers) combines global context (such as the structural stability of the entire mine face). After analysis, each sub-discriminator generates alternative discrimination results according to preset discrimination rules. For example, one sub-discriminator may output the preliminary conclusion that "the high-temperature area overlaps with structural cracks, posing a risk of collapse," while another sub-discriminator may output the conclusion that "the temperature distribution is uniform, but the local texture is abnormal, requiring inspection of equipment vibration." These alternative results are stored in a structured form, including the type of hazard, the location range, and the confidence score.

[0069] Many subtle details and textures in phosphate mine underground images are difficult to distinguish at low resolution. The high-resolution sub-discriminator D1 of the dual discriminator can better capture these details, resulting in clearer and more realistic images. The low-resolution sub-discriminator D2 effectively captures the overall lighting trend of the image, and combined with D1's detail discrimination, the generated image more closely matches the real underground scene in terms of lighting and texture. The dual discriminator, through a multi-scale discrimination method, enhances the sensitivity to lighting and details in phosphate mine underground images, solving the problem of single-scale discriminators easily overlooking details and lighting differences in low-light environments.

[0070] In this embodiment, when comprehensively processing the backup results output by the sub-discriminators, dynamic weights are first assigned based on the historical performance and current scenario adaptability of each sub-discriminator. For example, if a sub-discriminator has a high accuracy rate in similar past scenarios, its output result will be given a higher weight; if temperature characteristics are more critical in the scenario (such as fire warning), the weight of sub-discriminators that focus on temperature analysis will increase. Simultaneously, conflicts between backup results are detected (such as different sub-discriminators giving inconsistent hazard type judgments for the same area), and these conflicts are resolved through multimodal cross-validation. For example, if one sub-discriminator judges it as "structural softening caused by high temperature" and another judges it as "cracks caused by mechanical impact," the temperature distribution and vibration data are combined to verify which explanation is more reasonable, retaining the conclusion that best conforms to physical laws.

[0071] After weight allocation and conflict resolution, the backup results from each sub-discriminator are merged into the final hazard identification result. The fusion process employs a weighted voting mechanism, combining the confidence scores and dynamic weights of each sub-discriminator to calculate the overall hazard probability. For example, if two of the three sub-discriminators determine that the high-temperature area has a collapse risk (confidence levels of 0.8 and 0.7 respectively), and one determines that there is no significant hazard (confidence level of 0.3), the overall result will be classified as a collapse risk with a high probability. Furthermore, a redundancy verification mechanism is introduced. If the confidence level of the overall result is below a threshold, or if the results from multiple sub-discriminators differ significantly, a manual review process is triggered. The final hazard identification result is output in a structured report format, including the hazard type, location, severity, and recommended measures, ensuring the reliability and interpretability of the decision.

[0072] This application's technical solution acquires phosphate mine data using an intelligent video sensor; aligns the multimodal phosphate mine data to generate a spatiotemporally aligned data matrix; performs attention-based linear processing on the spatial dimensions corresponding to the data matrix to obtain a feature matrix; processes the feature matrix using a pre-defined dilated convolution and residual network to generate image features with consistent brightness; inputs the image features into a pre-defined number of sub-discriminators, each outputting backup discrimination results; and comprehensively processes these backup discrimination results to generate a hazard identification result. The intelligent video sensor acquires multimodal phosphate mine data, and the alignment processing ensures precise spatiotemporal correspondence, laying the foundation for accurate analysis. Attention-based linear processing focuses on key features, while dilated convolution and residual networks ensure consistent brightness, improving feature quality. Finally, the sub-discriminators comprehensively output hazard results, enabling multi-stage collaboration for comprehensive and accurate identification of underground phosphate mine hazards, ensuring production safety.

[0073] The following describes embodiments of the machine vision-based underground hazard identification system for phosphate mines according to this application, which can be used to execute the machine vision-based underground hazard identification method for phosphate mines described in the above embodiments of this application. It is understood that the machine vision-based underground hazard identification system for phosphate mines can be a computer program (including program code) running on a computer device; for example, the machine vision-based underground hazard identification system for phosphate mines can be an application software. This machine vision-based underground hazard identification system for phosphate mines can be used to execute the corresponding steps in the method provided in the embodiments of this application. For details not disclosed in the embodiments of the machine vision-based underground hazard identification system for phosphate mines of this application, please refer to the embodiments of the machine vision-based underground hazard identification method for phosphate mines described above.

[0074] H, W, C A block diagram of a machine vision-based underground hazard identification system for phosphate mines according to an embodiment of this application is shown.

[0075] Reference i, j, c As shown, a machine vision-based underground hazard identification system for phosphate mines according to an embodiment of this application includes: The acquisition module 310 acquires phosphate rock data through a smart video sensor; Alignment module 320 is used to align the multimodal phosphate rock data to generate a spatiotemporally aligned data matrix; Linear module 330 is used to perform attention-based linear processing on the spatial dimension corresponding to the data matrix to obtain the feature matrix; The extraction module 340 is used to process the feature matrix through a preset dilated convolution and residual network to generate image features with consistent brightness. The judgment module 350 is used to input the image features into a preset number of sub-discriminators, output backup judgment results respectively, and perform comprehensive processing on the backup judgment results to generate hidden danger judgment results.

[0076] In this application, based on the aforementioned scheme, the step of aligning the multimodal phosphate rock data to generate a spatiotemporally aligned data matrix includes: filtering the multimodal phosphate rock data based on the energy gradient of the phosphate rock data to generate first data; and performing multimodal data alignment based on dynamic time warping on the first data to generate a spatiotemporally aligned data matrix.

[0077] In this application, based on the aforementioned scheme, the step of performing attention-based linear processing on the spatial dimension corresponding to the data matrix to obtain the feature matrix includes: performing average pooling on the data matrix in the spatial dimension to obtain channel statistics; performing convolution and activation processing on the statistics to generate channel parameters; and performing linear processing on the channel dimension based on the channel parameters to obtain the attention-based feature matrix.

[0078] In this application, based on the aforementioned scheme, the step of performing average pooling on the data matrix in the spatial dimension to obtain channel statistics includes: performing average pooling on the data matrix in the spatial dimension to obtain channel statistics. c Statistical information for: in, Figure 3 These represent the height, width, and number of channels, respectively. Figure 3 These represent the height indicator, width indicator, and channel indicator, respectively.

[0079] In this application, based on the aforementioned scheme, the step of performing convolution and activation processing on the statistical information to generate channel parameters includes: performing convolution processing on the statistical information, followed by activation processing to generate channels. c Channel parameters for: in, Indicates the convolution kernel weights. Indicates channel c Statistical information H, W, C i, j, c ; These represent mapping processing and activation processing, respectively.

[0080] In this application, based on the aforementioned scheme, the step of processing the feature matrix through a preset dilated convolution and residual network to generate image features with consistent brightness includes: performing dilated convolution and activation processing on the feature matrix to generate activation features; and performing residual connections on the activation features based on the residual network to generate image features with consistent brightness.

[0081] In this application, based on the aforementioned scheme, the step of inputting the image features into a preset number of sub-discriminators, outputting backup discrimination results respectively, and comprehensively processing the backup discrimination results to generate a hazard discrimination result includes: inputting the image features into a preset number of sub-discriminators, outputting backup discrimination results respectively; and comprehensively processing the backup discrimination results based on the discrimination parameters preset for the sub-discriminators to generate a hazard discrimination result.

[0082] This application's technical solution acquires phosphate mine data using an intelligent video sensor; aligns the multimodal phosphate mine data to generate a spatiotemporally aligned data matrix; performs attention-based linear processing on the spatial dimensions corresponding to the data matrix to obtain a feature matrix; processes the feature matrix using a pre-defined dilated convolution and residual network to generate image features with consistent brightness; inputs the image features into a pre-defined number of sub-discriminators, each outputting backup discrimination results; and comprehensively processes these backup discrimination results to generate a hazard identification result. The intelligent video sensor acquires multimodal phosphate mine data, and the alignment processing ensures precise spatiotemporal correspondence, laying the foundation for accurate analysis. Attention-based linear processing focuses on key features, while dilated convolution and residual networks ensure consistent brightness, improving feature quality. Finally, the sub-discriminators comprehensively output hazard results, enabling multi-stage collaboration for comprehensive and accurate identification of underground phosphate mine hazards, ensuring production safety.

[0083] c and C represent channel identification and total number of channels, respectively Figure 4 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown.

[0084] It should be noted that the computer system of the electronic device in this embodiment is only an example and should not impose any limitations on the function and scope of use of the embodiments of this application.

[0085] In this embodiment, the computer system includes a central processing unit 401, which can perform various appropriate actions and processes based on a program stored in the read-only memory 402 or a program loaded from the storage section 408 into the random access memory 403, such as executing the machine vision-based method for identifying hidden dangers in phosphate mines described in the above embodiment. The random access memory 403 also stores various programs and data required for system operation. The central processing unit 401, the read-only memory 402, and the random access memory 403 are interconnected via a bus 404. An input / output interface 405 is also connected to the bus 404.

[0086] The following components are connected to the input / output interface 405: an input section 406 including a keyboard, mouse, etc.; an output section 407 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the input / output interface 405 as needed. A removable medium 411, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 410 as needed so that computer programs read from it can be installed into the storage section 408 as needed.

[0087] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 409, and / or installed from removable medium 411. When the computer program is executed by central processing unit 401, it performs various functions defined in the system of this application.

[0088] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0089] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0090] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0091] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various alternative implementations described above.

[0092] In another aspect, this application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to implement the machine vision-based method for identifying hidden dangers in underground phosphate mines as described in the above embodiments.

[0093] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0094] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the methods according to the embodiments of this application.

[0095] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.

[0096] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A method for identifying hidden dangers in underground phosphate mines based on machine vision, characterized in that, include: Acquire phosphate rock data using intelligent video sensors; The multimodal phosphate rock data is aligned to generate a spatiotemporally aligned data matrix; The spatial dimensions corresponding to the data matrix are subjected to attention-based linear processing to obtain the feature matrix; The feature matrix is ​​processed by a pre-defined dilated convolution and residual network to generate image features with consistent brightness. The image features are input into a preset number of sub-discriminators, which output backup discrimination results respectively. The backup discrimination results are then comprehensively processed to generate a hidden danger discrimination result.

2. The method for identifying hidden dangers in underground phosphate mines based on machine vision according to claim 1, characterized in that, Alignment processing is performed on the multimodal phosphate rock data to generate a spatiotemporally aligned data matrix, including: Based on the energy gradient of the phosphate rock data, the multimodal phosphate rock data is filtered to generate the first data. The first data is subjected to multimodal data alignment based on dynamic time warping to generate a spatiotemporally aligned data matrix.

3. The method for identifying hidden dangers in underground phosphate mines based on machine vision according to claim 1, characterized in that, The spatial dimensions corresponding to the data matrix are subjected to attention-based linear processing to obtain the feature matrix, including: The data matrix is ​​averaged in the spatial dimension to obtain the statistical information of the channels; The statistical information is subjected to convolution and activation processing to generate channel parameters; Based on the channel parameters, the channel dimensions are linearly processed to obtain an attention-based feature matrix.

4. The method for identifying hidden dangers in underground phosphate mines based on machine vision according to claim 3, characterized in that, The data matrix is ​​averaged in the spatial dimension to obtain channel statistics, including: The data matrix is ​​subjected to average pooling in the spatial dimension to obtain channels. c Statistical information for: in, H, W These represent the height and width of the channel, respectively. i, j, c These represent the height indicator, width indicator, and channel indicator, respectively.

5. The method for identifying hidden dangers in underground phosphate mines based on machine vision according to claim 4, characterized in that, The statistical information is subjected to convolution and activation processing to generate channel parameters, including: After performing convolution processing on the statistical information, activation processing is then performed to generate channels. c Channel parameters for: in, Indicates the convolution kernel weights, Indicates channel c Statistical information c and C These represent the channel identifier and the total number of channels, respectively. These represent mapping processing and activation processing, respectively.

6. The method for identifying hidden dangers in underground phosphate mines based on machine vision according to claim 1, characterized in that, The feature matrix is ​​processed using a pre-defined dilated convolution and residual network to generate image features with consistent brightness, including: The feature matrix is ​​subjected to dilated convolution and activation processing to generate activation features; The activation features are residually connected using a residual network to generate image features with consistent brightness.

7. The method for identifying hidden dangers in underground phosphate mines based on machine vision according to claim 1, characterized in that, The image features are input into a preset number of sub-discriminators, each outputting a backup discrimination result. These backup discrimination results are then comprehensively processed to generate a hazard discrimination result, including: The image features are input into a preset number of sub-discriminators, and alternative discrimination results are output respectively; Based on the preset discrimination parameters of the sub-discriminator, the backup discrimination results are comprehensively processed to generate hidden danger discrimination results.

8. A machine vision-based system for identifying hidden dangers in underground phosphate mines, characterized in that, include: The acquisition module acquires phosphate rock data through intelligent video sensors; The alignment module is used to align the multimodal phosphate rock data to generate a spatiotemporally aligned data matrix. The linear module is used to perform attention-based linear processing on the spatial dimensions corresponding to the data matrix to obtain the feature matrix; The extraction module is used to process the feature matrix through a preset dilated convolution and residual network to generate image features with consistent brightness. The judgment module is used to input the image features into a preset number of sub-discriminators, output backup judgment results respectively, and perform comprehensive processing on the backup judgment results to generate hidden danger judgment results.

9. The machine vision-based underground hazard identification system for phosphate mines according to claim 8, characterized in that, Alignment processing is performed on the multimodal phosphate rock data to generate a spatiotemporally aligned data matrix, including: Based on the energy gradient of the phosphate rock data, the multimodal phosphate rock data is filtered to generate the first data. The first data is subjected to multimodal data alignment based on dynamic time warping to generate a spatiotemporally aligned data matrix.

10. The machine vision-based underground hazard identification system for phosphate mines according to claim 8, characterized in that, The spatial dimensions corresponding to the data matrix are subjected to attention-based linear processing to obtain the feature matrix, including: The data matrix is ​​averaged in the spatial dimension to obtain the statistical information of the channels; The statistical information is subjected to convolution and activation processing to generate channel parameters; Based on the channel parameters, the channel dimensions are linearly processed to obtain an attention-based feature matrix.