Coal mine trackless rubber-tyred vehicle tire health state monitoring method based on multi-modal fusion

By using a multimodal fusion deep learning model to automatically assess the wear status of tires on trackless rubber-tired vehicles in coal mines, the problems of low efficiency and significant environmental impact of manual monitoring have been solved, achieving efficient and accurate wear monitoring.

CN122492576APending Publication Date: 2026-07-31CHINA COAL RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA COAL RES INST
Filing Date
2026-04-16
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In existing technologies, tire wear monitoring of trackless rubber-tired vehicles in coal mines relies on manual visual inspection, which is inefficient, yields inconsistent results, and is greatly affected by the environment, making it difficult to meet the needs of continuous operation and posing safety risks.

Method used

A deep learning method based on multimodal fusion is adopted. By acquiring visual images, acoustic signals and physical sensor data, and using a pre-trained multimodal fusion deep learning model, wear category, wear depth value and heat map are output. Threshold segmentation and connected component analysis are also performed to achieve automated wear status assessment.

Benefits of technology

It has achieved full automation of tire wear monitoring, avoiding the subjectivity and environmental limitations of manual monitoring, and significantly improving monitoring efficiency and positioning accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492576A_ABST
    Figure CN122492576A_ABST
Patent Text Reader

Abstract

This disclosure relates to a method for monitoring the health status of tires on trackless rubber-tired vehicles in coal mines based on multimodal fusion. The method includes: acquiring multimodal data of the tires of the trackless rubber-tired vehicles to be monitored; inputting the multimodal data into a pre-trained multimodal fusion deep learning model to obtain the tire wear category and wear depth value, as well as a heatmap corresponding to the wear category; performing threshold segmentation on the heatmap to obtain a binary image; performing connected component analysis on the binary image, and determining the pixel-level wear region in the visual image corresponding to the wear category based on the results of the connected component analysis. This disclosure improves the efficiency and positioning accuracy of tire health status monitoring for trackless rubber-tired vehicles in coal mines.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of coal mining technology, and in particular to a method for monitoring the health status of tires of trackless rubber-tired vehicles in coal mines based on multimodal fusion. Background Technology

[0002] In related technologies, trackless rubber-tired vehicles in coal mines are crucial equipment for underground material transportation and personnel transfer. Their tires, as key load-bearing components, operate under harsh and complex road conditions for extended periods, resulting in significant wear. Excessive wear significantly reduces vehicle traction, affecting driving safety and transportation efficiency, and in severe cases, may lead to major accidents such as tire blowouts. Currently, coal mine tire wear monitoring mainly relies on manual visual inspection, with monitoring personnel judging the degree of wear based on experience observing the tire surface condition. However, this traditional method has many shortcomings: low monitoring efficiency, making it difficult to meet the needs of continuous operation; results are greatly influenced by subjective factors, with inconsistent judgment standards among different monitoring personnel, easily leading to missed detections and misjudgments; and the insufficient lighting and dusty environment underground further increases the difficulty and safety risks of manual monitoring. Summary of the Invention

[0003] To overcome the problems existing in related technologies, this disclosure provides a method for monitoring the health status of tires of trackless rubber-tired vehicles in coal mines based on multimodal fusion.

[0004] According to a first aspect of the present disclosure, a method for monitoring the health status of tires on trackless rubber-tired vehicles in coal mines based on multimodal fusion is provided, comprising:

[0005] Acquire multimodal data of the tires of the trackless rubber-tired vehicle to be monitored in the coal mine; the multimodal data includes at least visual images; The multimodal data is input into a pre-trained multimodal fusion deep learning model to obtain the wear category and wear depth value of the tire, as well as a heat map corresponding to the wear category; wherein, the heat map is used to indicate the image region in the model used to distinguish the wear category; The heatmap is subjected to threshold segmentation to obtain a binary image; Perform connected component analysis on the binary image, and determine the pixel-level wear region in the visual image corresponding to the wear category based on the results of the connected component analysis.

[0006] According to a second aspect of the present disclosure, a tire health status monitoring device for trackless rubber-tired vehicles in coal mines based on multimodal fusion is provided, comprising: An acquisition unit is used to acquire multimodal data of the tires of the trackless rubber-tired vehicle to be monitored in the coal mine; the multimodal data includes at least visual images; The generation unit is used to input the multimodal data into a pre-trained multimodal fusion deep learning model to obtain the wear category and wear depth value of the tire, as well as a heat map corresponding to the wear category; wherein, the heat map is used to indicate the image region in the model used to distinguish the wear category; A segmentation unit is used to perform threshold segmentation processing on the heatmap to obtain a binary image; The determining unit is used to perform connected component analysis on the binary image and determine the pixel-level wear region in the visual image corresponding to the wear category based on the result of the connected component analysis.

[0007] According to a third aspect of the present disclosure, an electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method as described in any one of the first aspects.

[0008] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method as described in any one of the first aspects.

[0009] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method as described in any one of the first aspects.

[0010] The technical solutions provided by the embodiments of this disclosure can include the following beneficial effects: by acquiring multimodal data of the tire to be monitored and inputting it into a pre-trained multimodal fusion deep learning model, the wear category, wear depth value, and heat map indicating the discrimination area of ​​the tire can be output, realizing a multi-dimensional joint evaluation of the wear state; on this basis, the generated heat map is subjected to threshold segmentation and connected component analysis in sequence, and finally the pixel-level wear area corresponding to the wear category in the visual image is determined, thereby refining the image-level classification result into specific wear location information without the need for pixel-level annotation. This not only realizes the full-process automation of tire wear monitoring, effectively avoiding the subjectivity of manual monitoring and the limitations of the underground environment, but also completes wear level judgment, depth estimation, and area positioning simultaneously through an end-to-end deep learning model, significantly improving monitoring efficiency and positioning accuracy.

[0011] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0012] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0013] Figure 1 This is a flowchart illustrating a method for monitoring the health status of tires of trackless rubber-tired vehicles in coal mines based on multimodal fusion, according to an exemplary embodiment.

[0014] Figure 2 This is a block diagram illustrating a tire health status monitoring device for trackless rubber-tired vehicles in coal mines based on multimodal fusion, according to an exemplary embodiment.

[0015] Figure 3 This is a block diagram illustrating an apparatus for a method of monitoring the health status of tires of trackless rubber-tired vehicles in coal mines based on multimodal fusion, according to an exemplary embodiment.

[0016] Figure Labels 201 - Acquisition unit; 202 - Generation unit; 203 - Segmentation unit; 204 - Determination unit; 300 - Device; 302 - Processing component; 304 - Memory; 306 - Power component; 308 - Multimedia component; 310 - Audio component; 312 - I / O interface; 314 - Sensor component; 316 - Communication component; 320 - Processor. Detailed Implementation

[0017] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0018] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the present disclosure. The singular forms “a” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.

[0019] It should be understood that although the terms first, second, third, etc., may be used to describe various information in embodiments of this disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first information may also be referred to as second information without departing from the scope of embodiments of this disclosure, and similarly, second information may also be referred to as first information. Depending on the context, the words “if” and “suppose” as used herein may be interpreted as “when”, “when”, or “in response to a determination”.

[0020] Furthermore, various forms of processes shown in the embodiments of this disclosure can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0021] It should be noted that the collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0022] Figure 1 This is a flowchart illustrating a method for monitoring the tire health status of a trackless rubber-tired vehicle in a coal mine based on multimodal fusion, according to an exemplary embodiment. Figure 1 As shown, it should be noted that the multimodal fusion-based method for monitoring the tire health status of trackless rubber-tired vehicles in coal mines, as described in this disclosure, is applied to a multimodal fusion-based tire health status monitoring device for trackless rubber-tired vehicles in coal mines. Figure 1 As shown, the method may include the following steps: In some embodiments of this disclosure, a multimodal fusion deep learning model needs to be pre-trained before acquiring multimodal data of the tires of the trackless rubber-tired vehicles to be monitored in the coal mine. Specifically, the training process includes the following steps: Step a1: Obtain multimodal data samples of the tires of the trackless rubber-tired vehicle in the coal mine, and label each sample with an image-level wear category label and a wear depth numerical label.

[0023] As an example, visual images, acoustic signals, and physical sensor data can be collected during vehicle entry, exit, or maintenance using explosion-proof high-definition cameras, high-sensitivity microphones, and tire pressure and acceleration sensors. Then, professional monitoring personnel label each image with an image-level wear category (e.g., "no wear," "light wear," "moderate wear," "heavy wear"), while simultaneously using a high-precision laser profilometer to measure the actual wear depth as a numerical label.

[0024] In some embodiments, data augmentation techniques such as rotation, flipping, cropping, brightness adjustment, and Gaussian blur can be used to expand the dataset and improve the model's generalization ability and robustness.

[0025] Understandably, using image-level labels instead of pixel-level segmentation labels can significantly reduce annotation costs while meeting the training requirements for weakly supervised localization.

[0026] Step a2: Preprocess the multimodal data samples to obtain aligned multimodal training features.

[0027] Specifically, the visual image is sequentially subjected to median filtering for denoising, CLAHE image enhancement, and YOLOv8-based tire region cropping to obtain a standardized target tire image. The acoustic signal is then bandpass filtered for denoising and converted to a Mel spectrogram via short-time Fourier transform. Time-frequency analysis is performed on the vibration / acceleration data to extract frequency domain features, and tire pressure / temperature data is normalized to obtain scalar features. These two are then combined to form the target features. Next, using the global timing system of the edge computing gateway as a benchmark, sliding window interpolation and resampling methods are employed to synchronize and align the target tire image, Mel spectrogram, and target features at the millisecond level. Finally, the aligned multimodal data is feature-stitched to obtain multimodal training features. This preprocessing step effectively eliminates the influence of downhole environmental noise and uneven illumination, providing high-quality, well-aligned multimodal input for subsequent model training.

[0028] In some embodiments, the original image is first processed using median filtering. This filtering method replaces the center pixel value with the median of the neighboring pixel values, effectively removing salt-and-pepper noise caused by coal mine dust particles while preserving the edge information of the tire tread pattern. Then, the CLAHE (Contrast-Limited Adaptive Histogram Equalization) algorithm is applied to the denoised image. The image is divided into multiple small blocks, and histogram equalization is performed on each block, limiting the contrast amplification. This significantly improves the contrast and visibility of tire wear texture in low-light, high-dust environments without over-enhancing noise. Finally, a pre-trained YOLOv8 target detection model is used to automatically locate and crop the tire region, removing background interference so that subsequent feature extraction focuses on key areas of the tire tread and sidewall. This preprocessing effectively overcomes the impact of the complex underground environment on image quality, providing standardized, high-quality visual input for subsequent multimodal fusion models.

[0029] Step a3: Construct an initial multimodal fusion deep learning model, which includes a visual sub-model, an acoustic sub-model, a physical sub-model, a multimodal fusion module, a classification output head, a regression output head, and a weakly supervised localization output head.

[0030] As an example, the visual sub-model employs a Swin Transformer backbone network to extract multi-scale features from visual images through a self-attention mechanism of the moving window; the acoustic sub-model uses a lightweight convolutional neural network to extract acoustic features from the Mel spectrogram; the physical sub-model uses a fully connected network to extract physical features from physical sensor data; the multimodal fusion module concatenates the features from the three branches along the channel dimension and then reduces the dimensionality through a fully connected layer to obtain the fused features; the classification output head, regression output head, and weakly supervised localization output head are connected after the fused features to output wear category, wear depth value, and heatmap, respectively. This model architecture can fully utilize the complementarity of multimodal information to achieve end-to-end multi-task learning.

[0031] Step a4: Input the multimodal training features into the initial multimodal fusion deep learning model to obtain the wear category prediction value, wear depth prediction value, and heat map prediction value.

[0032] In one embodiment, the multimodal training features generated in step a1 are fed into the model in batches for forward propagation. The visual, acoustic, and physical branches extract deep features of their respective modalities, which are then fused and used by the three output heads to generate prediction results.

[0033] Among them, the weakly supervised localization output head uses the Grad-CAM (Gradient-weighted Class Activation Mapping) algorithm to generate heatmap prediction values.

[0034] Understandably, this step enables the model to reason about the input data, providing a foundation for subsequent loss calculations.

[0035] Step a5: Calculate the classification loss between the image-level wear category label and the wear category prediction value; calculate the regression loss between the wear depth numerical label and the wear depth prediction value; calculate the attention consistency loss between the heatmap prediction value and the acoustic attention map.

[0036] Specifically, the classification loss uses the cross-entropy loss function, and the regression loss uses the smoothed L1 loss function. The attention consistency loss is used to force the model to maintain spatial consistency between the visual heatmap and the acoustic attention map, where the acoustic attention map is an attention map corresponding to the wear category generated by class activation mapping based on the feature maps of the acoustic sub-models in the multimodal fusion deep learning model.

[0037] As an example, the attention consistency loss uses cosine similarity to measure the difference between the two. By introducing this loss, it can be ensured that the anomalous frequencies captured by the acoustic features do indeed originate from the wear points located visually, thereby eliminating false alarms caused by background noise.

[0038] Step a6: The model parameters are updated by backpropagation using the weighted sum of attention consistency loss, classification loss, and regression loss as the total loss, so that the initial multimodal fusion deep learning model can simultaneously learn the wear classification task, wear depth regression task, and heatmap generation task.

[0039] In one embodiment, the formula for calculating the total loss is:

[0040] in, These are weighting coefficients (e.g., λ1=1.0, λ2=0.5, λ3=0.2), used to constrain different modes to focus on the same physical wear region. For classifying losses, To regress the loss, This is an attention consistency loss. The AdamW optimizer, combined with cosine annealing learning rate scheduling, can be used to iteratively update the model parameters on the training set until convergence.

[0041] In some embodiments of this disclosure, the wear category classification loss can employ the cross-entropy loss function to measure the model's accuracy in classifying categories such as "no wear," "light wear," "moderate wear," and "heavy wear." In one embodiment, the classification loss... The calculation formula can be expressed as:

[0042] in, It is the number of categorical samples. It is the number of categories. It's a real label. It is a predicted probability.

[0043] The wear depth regression loss, typically using Smooth L1 Loss or Mean Squared Error (MSE), is used to measure the accuracy of the model's numerical prediction of a specific wear depth.

[0044]

[0045] in The number of samples with regression labels. It is the actual wear depth. It predicts the depth of wear. The function is defined as:

[0046] The attention consistency loss, designed to encourage the model to focus on real wear areas, forces the model to pay attention to the visual heatmap. Acoustic feature map Maintaining spatial consistency in the mapping ensures that anomalous frequencies captured by acoustic features do indeed originate from visually located wear points, thereby eliminating false alarms caused by background noise.

[0047] In some embodiments of this disclosure, the AdamW optimizer can be selected as the parameter update algorithm. This optimizer introduces weight decay decoupling on the basis of Adam, which can effectively suppress overfitting and is particularly suitable for training multi-task deep learning models. In one embodiment, the initial learning rate is set to 1e-4, and the weight decay coefficient is set to 1e-2. At the same time, in conjunction with the cosine annealing learning rate scheduling strategy, the learning rate is dynamically adjusted according to the cosine function curve in each training cycle: the learning rate decreases slowly in the initial stage to stabilize exploration, decreases rapidly in the middle stage to achieve fast convergence, and is further reduced to near zero in the final stage to finely adjust the model parameters.

[0048] As an example, the total number of training epochs is set to 100, and the learning rate gradually decreases from 1e-4 to 1e-6. In each iteration, the total loss L for the current batch is calculated. total The gradient is calculated via backpropagation, and then the parameters are updated by the AdamW adaptive moment estimator optimizer with weighted decay. This process is repeated until the loss value on the training set converges (e.g., after 10 consecutive epochs verifying that the loss no longer decreases). This optimized combination achieves better model generalization performance while maintaining convergence speed, thereby improving the accuracy of tire wear classification, deep regression, and heatmap generation tasks.

[0049] In some embodiments of this disclosure, the classification loss may also employ a binary cross-entropy loss function to measure the model's accuracy in distinguishing between the "worn" and "unworn" states. Specifically, for a binary classification task, the classification loss L... BCE The calculation formula can be expressed as:

[0050] in, It is the sample size. It is a real label (0 or 1). It is the probability predicted by the model.

[0051] Furthermore, data augmentation techniques such as random rotation, brightness adjustment, and spectral masking can be employed, along with an early stopping mechanism to prevent overfitting. This training strategy enables the model to learn pixel-level localization capabilities solely based on image-level labels, significantly reducing reliance on expensive pixel-level annotations.

[0052] Once training is complete, the model can be used to monitor the real-time health status of tires on trackless rubber-tired vehicles in coal mines. The specific steps include: Step 101: Obtain multimodal data of the tires of the trackless rubber-tired vehicle to be monitored in the coal mine; the multimodal data includes at least visual images.

[0053] In some embodiments of this disclosure, the visual stream can acquire clear images of key areas such as tire tread, sidewall, and tread texture by deploying explosion-proof high-definition cameras, ensuring high-quality visual data acquisition even in complex environments. The cameras should be fixed at multiple angles or used in conjunction with an automatic rotating bracket to ensure clear images of key wear areas such as the tire tread and sidewall.

[0054] Multimodal data can also include acoustic streams, which are collected by deploying high-sensitivity microphones to capture the acoustic signature of tire-ground friction during vehicle operation, and used to analyze tire wear and abnormal sounds.

[0055] Multimodal data can also include physical flow, using built-in tire pressure and acceleration sensors to monitor tire pressure, tire temperature and wheel vibration frequency in real time, and obtain data on physical changes during vehicle operation to help determine tire wear and safety.

[0056] This can be done during vehicle entry, exit, routine parking, or maintenance, ensuring the vehicle is stationary or at low speed.

[0057] In some embodiments of this disclosure, the multimodal data further includes acoustic signals and physical sensing data. Therefore, step 101 may specifically include the following sub-steps: Step b1 involves denoising, enhancing, and cropping the tire region of the visual image to obtain a standardized target tire image.

[0058] Specifically, median filtering is used to remove noise caused by dust particles, the CLAHE algorithm is used to limit contrast enhancement, and finally, a pre-trained target detection model (YOLOv8) is used to locate and crop the tire area. This processing can eliminate the interference of low lighting and dust in the well to the image quality and highlight the tire wear texture.

[0059] Furthermore, the enhanced image at the pixel level is calculated using the following formula. grayscale value at :

[0060] in, The original image at the pixel level The grayscale value at that location. Therefore The local neighborhood centered on the center. It is based on local neighborhood The grayscale transformation function is calculated from the histogram (with its cropping limits).

[0061] Step b2 involves denoising the acoustic signal and converting the denoised acoustic signal into a Mel spectrogram.

[0062] Specifically, a bandpass filter (e.g., retaining the 500Hz-4000Hz frequency band) is used to remove low-frequency environmental noise such as that from underground ventilation fans. Then, a short-time Fourier transform is used to convert the one-dimensional sound wave into a two-dimensional Mel-spectrum image, which facilitates spatial alignment with visual features. This transformation allows voiceprint features to be input into a deep learning model in image form, achieving unified cross-modal processing.

[0063] Step b3: Perform time-frequency analysis on the vibration data and / or acceleration data in the physical sensing data to obtain frequency domain features; normalize the tire pressure data and / or tire temperature data in the physical sensing data to obtain normalized scalar features; combine the frequency domain features and scalar features to obtain the target features.

[0064] As an example, a short-time Fourier transform is performed on the vibration signal to extract the frequency domain energy distribution. Tire pressure and tire temperature are normalized and mapped to the [0,1] interval, and then the two are concatenated into a one-dimensional feature vector. This step can extract complementary information reflecting the dynamic load and inflation state of the tire from the physical signal.

[0065] Step b4: Synchronize and align the target tire image, Mel spectrogram, and target features in time to obtain aligned multimodal data.

[0066] Specifically, using the timestamp of each frame as the center, audio slices of corresponding duration are extracted from the acoustic stream and a Mel spectrogram is generated. Simultaneously, spline interpolation is used to resample the physical feature sequence to ensure strict temporal alignment of the data from the three modalities. This alignment operation is fundamental to multimodal fusion, guaranteeing that information from different sources reflects the tire state at the same moment.

[0067] Step b5: Perform feature fusion on the aligned multimodal data to obtain feature-fused multimodal data.

[0068] In one embodiment, the target tire image, Mel spectrogram, and target features are directly concatenated along the channel dimension to form a high-dimensional feature tensor. This fusion method is simple and efficient, preserving the original information of each modality for subsequent deep learning models to extract deeper features.

[0069] Through steps a1 to a5 above, high-quality, aligned, and fused multimodal data can be obtained. Understandably, this preprocessing workflow effectively overcomes the vulnerability of single-modal data in coal mines to environmental interference, providing reliable input for model inference.

[0070] Step 102: Input the multimodal data into the pre-trained multimodal fusion deep learning model to obtain the tire wear category and wear depth value, as well as the heat map corresponding to the wear category.

[0071] The heatmap is used to indicate the image regions in the model used to determine the wear category.

[0072] In some embodiments of this disclosure, multimodal data is input into a pre-trained multimodal fusion deep learning model. The multimodal fusion deep learning model includes a visual sub-model, an acoustic sub-model, a physical sub-model, a multimodal fusion module, a classification output head, a regression output head, and a weakly supervised localization output head, wherein: The visual sub-model employs a Swing Transformer backbone network to extract multi-scale features from visual images through a self-attention mechanism using a moving window. The visual sub-model includes a channel attention module to weight different feature channels, enhancing sensitivity to tire wear texture features, and a spatial attention module to focus on key spatial locations of wear, improving the localization accuracy of tire wear texture features.

[0073] As an example, the Swin Transformer divides the image into non-overlapping windows and computes self-attention within each window. Cross-window information interaction is achieved by moving the window, thus obtaining a global receptive field with linear computational complexity. The channel attention module employs a squeeze-and-excitation structure to automatically learn the weights of each channel; the spatial attention module uses the spatial branch from the convolutional block attention module to generate a spatial attention map. This design enables the model to simultaneously focus on important feature channels and important spatial locations, significantly improving the representation ability of wear features.

[0074] Furthermore, the squeeze operation – Global Average Pooling (GAP) can specifically include the following steps: Compress the spatial information of each channel into a global descriptor:

[0075] in, It is the first of the input feature maps One channel (dimension is) ), Is this channel at the pixel? The value at that location, It is the first The global average pooling result for each channel.

[0076] Excitation operation - channel weight generation: Learns the non-linear relationship between channels through two fully connected layers (or 1x1 convolutions) and outputs the weight of each channel:

[0077] in, It is made by all The vector formed It is the learned weight matrix (corresponding to the parameters of the fully connected layer). It is the ReLU activation function. It is the Sigmoid activation function, which normalizes the weight values ​​to the range (0, 1). This is the final weight vector generated for each channel, with each element... Representing the The importance of each channel.

[0078] Scale operation - feature recalibration: Multiply the learned weights back onto each channel of the original feature map:

[0079] in, It is the first one after channel attention weighting. Each feature channel.

[0080] An acoustic sub-model is used to extract acoustic features from the Mel spectrogram.

[0081] In one embodiment, the acoustic submodel employs a lightweight convolutional neural network comprising two convolutional layers and one global average pooling layer, outputting a fixed-dimensional acoustic feature vector. This submodel is able to capture the frequency domain energy distribution of tire friction sound patterns from the Mel spectrogram, providing auditory evidence for wear assessment.

[0082] The physical sub-model is used to extract physical features from physical sensing data.

[0083] As an example, the physical sub-model employs a three-layer fully connected network, taking normalized tire pressure, tire temperature, and vibration frequency domain features as inputs, and outputting a physical feature vector. This sub-model reflects the tire's load, temperature, and dynamic impact state in real time, assisting in the identification of abnormal wear.

[0084] The multimodal fusion module is used to fuse visual features, acoustic features, and physical features to obtain fused features.

[0085] Specifically, the feature vectors output by the three sub-models are concatenated along the channel dimension and then reduced to 256 dimensions through a fully connected layer to form a fused feature. This fused feature integrates multi-dimensional information from visual texture, acoustic frequency, and physical parameters, providing a decision-making basis for the subsequent output head.

[0086] The classification output head is used to output the wear category of the tire based on the fused features.

[0087] In one embodiment, the classification output head consists of a fully connected layer and a Softmax activation function, outputting the probability distribution of four categories (no wear, mild, moderate, and severe).

[0088] The regression output head is used to output the tire wear depth value based on the fused features.

[0089] As an example, the regression output head uses a fully connected layer followed by a linear activation function to directly output the predicted wear depth value (in millimeters).

[0090] The weakly supervised localization output head is used to generate heatmaps corresponding to wear categories based on category activation maps or their variants, so as to complete pixel-level wear area localization based on image-level labels.

[0091] Specifically, the output head reuses the gradient information from the classification output head to generate a heatmap using the Grad-CAM algorithm. The value of each pixel in the heatmap represents the degree to which that location contributes to the model's determination of the current wear category.

[0092] It should be noted that the weakly supervised localization output head in this embodiment generates a heatmap using the following steps: extracting feature maps from the preset target convolutional layer in the multimodal fusion deep learning model; and calculating the heatmap corresponding to the wear category using the Grad-CAM formula.

[0093]

[0094] in, Indicates category (For example, in the Grad-CAM heatmap generated for "moderate wear"), higher brightness in the heatmap indicates that the area is classified as a certain category by the model. The greater the contribution. The last convolutional layer (or a specified layer) of a deep convolutional neural network Each feature map It is the model prediction category The score. The ReLU rectified linear unit activation function ensures that only positive correlations are retained, while removing negative correlation noise. Indicate category Relative to feature map Importance weights are assigned to the categories in the model output. Fractions of feature maps The gradient is obtained by taking the derivative and then performing global average pooling.

[0095] This formula can backpropagate deep semantic information to the input space, intuitively display the model's region of interest, and achieve pixel-level localization under weak supervision.

[0096] The multimodal fusion deep learning model of this disclosure also includes a multimodal feature refinement module, which adopts an SE channel attention mechanism and is set after the multimodal fusion module. It is used to recalibrate the cross-modal weights of the fused features to improve the weights of modal features that contribute highly to the current monitoring task.

[0097] Specifically, the refining module first compresses the spatial dimension of the fused features through global average pooling, then generates channel weight vectors through two fully connected layers and Sigmoid activation, and finally multiplies the weights back into the original fused features. Global attention from the SwingTransformer completes feature extraction within a single modality, and combined with the fused SE refiner, it performs weight redistribution across modalities, forming a hierarchical focusing system from "local details" to global semantics, and then to modal decision-making, significantly improving the model's discrimination accuracy in complex environments.

[0098] Step 103: Perform threshold segmentation on the heatmap to obtain a binary image.

[0099] In some embodiments of this disclosure, step 103 may specifically include the following sub-steps: Step c1: An adaptive threshold segmentation algorithm is used to calculate the segmentation threshold based on the distribution characteristics of pixel values ​​in the heatmap.

[0100] As an example, the Otsu algorithm can be used to calculate the optimal threshold that maximizes the inter-class variance between the foreground (high-response area) and the background in the heatmap.

[0101] Step c2 involves setting pixels in the heatmap with values ​​greater than or equal to the segmentation threshold to a first value (e.g., 255), and pixels with values ​​less than the segmentation threshold to a second value (e.g., 0), generating a binary image. In this binary image, white areas represent the set of pixels that the model considers to contribute significantly to wear detection, i.e., candidate wear regions. Thresholding segmentation transforms the continuous thermal response into a defined region mask, providing input for subsequent connected component analysis.

[0102] Step 104: Perform connected component analysis on the binary image, and determine the pixel-level wear regions in the visual image corresponding to the wear category based on the results of the connected component analysis.

[0103] In this embodiment, morphological operations are performed on the binary image. For example, erosion is first used to remove isolated noise points, and then dilation is used to connect regions broken due to discontinuities in the wear texture to smooth the wear area contour. Then, connected component analysis is performed on the processed binary image to identify all connected components and statistically analyze the area, centroid, and other attributes of each component. Finally, a minimum bounding rectangle or irregular contour is generated for each connected component with an area greater than a preset threshold, serving as the pixel-level wear area location result. Through the above processing, the heatmap can be transformed into intuitive and quantifiable pixel-level wear location information, providing accurate spatial basis for subsequent depth calibration and alarm decisions.

[0104] In some embodiments of this disclosure, step 104 may specifically include the following sub-steps: Step d1 involves performing morphological operations on the binary image to remove noise and broken connection regions; morphological operations include erosion and / or dilation.

[0105] Specifically, an etching operation is first performed to eliminate isolated noise points, followed by an expansion operation to connect adjacent fracture areas. This operation smooths the contour of the worn area and improves positioning accuracy.

[0106] Step d2 involves performing connected component analysis on the binary image after morphological operations to identify at least one connected component.

[0107] As an example, the eight-neighbor connectivity algorithm is used to label all connected regions and calculate the area of ​​each region.

[0108] Step d3 generates a minimum bounding rectangle or irregular contour for each connected component to obtain the pixel-level wear region localization result. For connected regions with an area larger than a preset threshold (e.g., 50 pixels), its minimum bounding rectangle is generated; for scenes requiring fine shapes, irregular contours can be generated by finding the convex hull or edge tracking algorithms. This localization result can be directly overlaid on the original tire image to intuitively indicate the location of wear.

[0109] In some embodiments of this disclosure, after determining the pixel-level wear region in the visual image corresponding to the wear category based on the results of connected component analysis, geometric features and texture features are extracted from the pixel-level wear region, and the geometric features and texture features are input into a multimodal fusion deep learning model to calibrate the wear depth value of the pixel-level wear region.

[0110] In this embodiment of the disclosure, geometric features include at least one of area, perimeter and aspect ratio, and texture features include at least one of gray-level co-occurrence matrix (GLCM) features (such as contrast, energy, homogeneity) and local binary mode (LBP) features.

[0111] Specifically, within the wear region obtained in step d3, the pixel area, contour perimeter, and aspect ratio of the minimum bounding rectangle of the region are calculated. Simultaneously, the GLCM and LBP histograms of the pixels within this region are calculated to form a texture feature vector. This feature vector, along with the wear depth prediction value output in step b6, is then input into a lightweight calibration network (e.g., a two-layer fully connected network) to output the calibrated wear depth value.

[0112] As an example, the calibration network is jointly trained with the main model using a multi-task learning approach, utilizing micro-texture information within the wear area to correct the initial depth prediction.

[0113] This step can further improve the accuracy of wear depth estimation, especially in the wear edge area or under poor lighting conditions.

[0114] In some embodiments of this disclosure, the method further includes the following two parallel steps: Step e1: Obtain the entropy value of the visual feature map output by the visual sub-model. When the entropy value is lower than the preset threshold, reduce the decision weight of the visual sub-model through the cross-attention module, while increasing the decision weight of the acoustic sub-model.

[0115] Specifically, the entropy value of each channel in the visual feature map is calculated and averaged. When the average entropy is less than a threshold (e.g., 0.3), it indicates that the visual features contain low information content (possibly due to insufficient lighting in the well or the lens being obscured by dust). In this case, the cross-attention module dynamically adjusts the weight coefficients of the three sub-models in the multimodal fusion module, reducing the contribution of the visual branch and increasing the contribution of the acoustic branch, utilizing the frequency shift characteristics of tire friction sound patterns to maintain monitoring accuracy. This compensation mechanism enables the system to operate reliably even in harsh visual environments.

[0116] Step e2: When physical sensor data is detected to be missing, the model is switched to the visual-acoustic dual-modal working mode, and the current wear depth is verified for consistency using the wear depth change trend in historical monitoring data.

[0117] For example, if the tire pressure sensor malfunctions, the model automatically ignores the input from the physical sub-model and infers solely based on visual and acoustic features. Simultaneously, the system reads the wear depth values ​​from the vehicle's past 10 monitoring sessions, fits a linear trend line, and compares the currently predicted wear depth value with the expected value on the trend line. If the deviation exceeds 20%, an alarm is triggered, indicating a sensor malfunction. This verification step ensures that the system maintains basic safety monitoring capabilities even when some sensors fail, preventing missed detections due to missing data.

[0118] In some embodiments of this disclosure, the method further includes the following sub-steps for identifying tire wear that is difficult to detect with the naked eye: Step f1: Extract acoustic signature spectral features from the acoustic signals in the multimodal data.

[0119] Specifically, a fast Fourier transform is performed on the original acoustic signal to obtain the power spectral density distribution.

[0120] Step f2: Monitor the energy density changes in the high-frequency bands related to the tire air pumping effect in the acoustic spectrum.

[0121] The air pumping effect refers to the high-frequency sound waves generated when air is compressed and released as the tire tread grooves come into contact with the ground. Its frequency range is typically between 1000Hz and 3000Hz. As the tire tread wears down, this high-frequency energy gradually weakens.

[0122] Step f3: When a nonlinear shift of the dominant frequency energy of the acoustic spectrum to a lower frequency band (e.g., 500Hz-800Hz) is detected, it is determined that the tire has experienced internal structural health degradation.

[0123] As an example, the system tracks the peak frequency of the spectrum in real time. If the peak frequency drops by more than 200Hz in a short period of time, it is determined to be an internal structural degradation.

[0124] Step f4: Calculate the audio mutual information entropy of the acoustic signature spectrum to convert the nonlinear offset feature into a wear degree quantification index value, and use the wear degree quantification index value to determine the hidden wear of the tires of the trackless rubber-tired vehicle in the coal mine.

[0125] Specifically, the difference between the current spectrum and the standard healthy spectrum is measured by mutual information entropy, and a quantitative value between 0 and 1 is output, with a larger value indicating more severe wear and tear. This indicator can be used independently as a warning basis, or it can be integrated with visual monitoring results to improve the reliability of the overall diagnosis.

[0126] In some embodiments of this disclosure, after determining the pixel-level wear area, wear type, and wear depth, the following operations can be performed: The wear area boundary box or outline is overlaid on the original tire image, and the wear type (e.g., "moderate wear") and wear depth (e.g., "2.5mm") are labeled next to the box. Simultaneously, a three-level alarm is set according to the wear degree: For light wear, the system records and prompts for attention during the next maintenance; for moderate wear, an orange warning is triggered, and the maintenance manager is notified via SMS; for heavy wear, a red alarm is triggered, activating an audible and visual alarm and notifying the vehicle to immediately stop operation via APP / telephone. Furthermore, the system automatically records the time, location, vehicle number, tire position, wear type, wear degree, and image evidence for each monitoring session, generating a historical monitoring report for data traceability and predictive maintenance. This step ensures that monitoring results can be promptly and effectively translated into safety actions.

[0127] According to the multimodal fusion-based method for monitoring the health status of tires of trackless rubber-tired vehicles in coal mines proposed in this disclosure, by acquiring multimodal data of the tire to be monitored and inputting it into a pre-trained multimodal fusion deep learning model, the method can output the tire wear category, wear depth value, and heat map indicating the discrimination region, thus realizing a multi-dimensional joint assessment of wear status. On this basis, the generated heat map is subjected to threshold segmentation and connected component analysis in sequence, and finally the pixel-level wear region corresponding to the wear category in the visual image is determined. Thus, without the need for pixel-level annotation, the image-level classification result is refined into specific wear location information. This not only realizes the full-process automation of tire wear monitoring, effectively avoiding the subjectivity of manual monitoring and the limitations of the underground environment, but also completes wear level judgment, depth estimation, and region positioning simultaneously through an end-to-end deep learning model, significantly improving monitoring efficiency and positioning accuracy.

[0128] Figure 2 This is a block diagram illustrating a multimodal fusion-based tire health status monitoring device for trackless rubber-tired vehicles in coal mines, according to an exemplary embodiment. (Refer to...) Figure 2 The device includes an acquisition unit 201, a generation unit 202, a segmentation unit 203, and a determination unit 204.

[0129] The acquisition unit 201 is used to acquire multimodal data of the tires of the trackless rubber-tired vehicle to be monitored in the coal mine; the multimodal data includes at least visual images. The generation unit 202 is used to input multimodal data into a pre-trained multimodal fusion deep learning model to obtain the wear category and wear depth value of the tire, as well as the heat map corresponding to the wear category; wherein, the heat map is used to indicate the image region in the model used to distinguish the wear category; The segmentation unit 203 is used to perform threshold segmentation processing on the heat map to obtain a binary image; The determination unit 204 is used to perform connected component analysis on the binary image and determine the pixel-level wear region in the visual image corresponding to the wear category based on the results of the connected component analysis.

[0130] In some embodiments of this disclosure, the multimodal data further includes acoustic signals and physical sensing data, and the apparatus further includes a preprocessing unit for: The visual image is denoised, enhanced, and the tire region is cropped to obtain a standardized target tire image. The acoustic signal is denoised, and the denoised acoustic signal is converted into a Mel spectrogram. Time-frequency analysis is performed on vibration and / or acceleration data from physical sensing data to obtain frequency domain features; tire pressure and / or tire temperature data from physical sensing data are normalized to obtain normalized scalar features; the frequency domain features and scalar features are combined to obtain the target features; The target tire image, Mel spectrogram, and target features are synchronized and aligned in time to obtain aligned multimodal data; Feature fusion is performed on the aligned multimodal data to obtain feature-fused multimodal data; The generation unit 202 can be specifically used to: input the multimodal data after feature fusion into a pre-trained multimodal fusion deep learning model to obtain the wear category and wear depth value of the tire, as well as the heat map corresponding to the wear category.

[0131] In some embodiments of this disclosure, the multimodal fusion deep learning model includes: The visual sub-model uses the Swin Transformer backbone network to extract multi-scale features of visual images through the self-attention mechanism of the moving window. The visual sub-model includes a channel attention module, which is used to weight different feature channels to enhance the sensitivity to tire wear texture features, and a spatial attention module, which is used to focus on key spatial locations of wear to enhance the localization accuracy of tire wear texture features. Acoustic sub-models are used to extract acoustic features from Mel spectrograms; The physical sub-model is used to extract physical features from physical sensing data. The multimodal fusion module is used to fuse visual features, acoustic features, and physical features to obtain fused features; A classification output head is used to output the wear category of the tire based on the fused features; The regression output head is used to output the tire wear depth value based on the fused features; The weakly supervised localization output head is used to generate heatmaps corresponding to wear categories based on category activation maps or their variants, so as to complete pixel-level wear area localization based on image-level labels.

[0132] In some embodiments of this disclosure, the generation unit 202 may specifically be used for: Extract feature maps from the pre-defined target convolutional layer in a multimodal fusion deep learning model; The heat map corresponding to the wear category is calculated using the following formula:

[0133]

[0134] in, Indicates the type of wear The generated Grad-CAM heatmap, The first target convolutional layer in a deep convolutional neural network Each feature map Wear category Relative to feature map Importance weight, Wear category The score.

[0135] In some embodiments of this disclosure, the segmentation unit 203 may specifically be used for: An adaptive threshold segmentation algorithm is used to calculate the segmentation threshold based on the distribution characteristics of pixel values ​​in the heatmap; Pixels with values ​​greater than or equal to the segmentation threshold in the heatmap are set as the first value, and pixels with values ​​less than the segmentation threshold are set as the second value, thus generating a binary image.

[0136] In some embodiments of this disclosure, the determining unit 204 may specifically be used for: Morphological operations are performed on binary images to remove noise and broken connectivity regions; morphological operations include erosion and / or dilation. Connectivity analysis is performed on the binary image after morphological operations to identify at least one connected component. For each connected component, a minimum bounding rectangle or irregular contour is generated to obtain the pixel-level location result of the worn area.

[0137] In some embodiments of this disclosure, the apparatus may further include a calibration unit for: Extract geometric and texture features from pixel-level wear regions; geometric features include at least one of area, perimeter and aspect ratio, and texture features include at least one of gray-level co-occurrence matrix features and local binary pattern features. Geometric and texture features are input into a multimodal fusion deep learning model to calibrate the wear depth values ​​of pixel-level wear regions.

[0138] In some embodiments of this disclosure, the apparatus may further include a training unit for: Obtain multimodal data samples of tires for trackless rubber-tired vehicles in coal mines, and label each sample with image-level wear category labels and wear depth numerical labels; Preprocess the multimodal data samples to obtain aligned multimodal training features; Construct an initial multimodal fusion deep learning model, which includes a visual sub-model, an acoustic sub-model, a physical sub-model, a multimodal fusion module, a classification output head, a regression output head, and a weakly supervised localization output head; The multimodal training features are input into the initial multimodal fusion deep learning model to obtain wear category predictions, wear depth predictions, and heat map predictions. Calculate the classification loss between the image-level wear category label and the wear category prediction value, and calculate the regression loss between the wear depth numerical label and the wear depth prediction value; Calculate the attention consistency loss between the heatmap prediction and the acoustic attention map; the acoustic attention map is an attention map corresponding to the wear category generated by class activation mapping based on the feature map of the acoustic sub-model in the multimodal fusion deep learning model. By using the weighted sum of attention consistency loss, classification loss, and regression loss as the total loss, the model parameters are updated through backpropagation, enabling the initial multimodal fusion deep learning model to simultaneously learn wear classification tasks, wear depth regression tasks, and heatmap generation tasks.

[0139] In some embodiments of this disclosure, the apparatus may further include a verification unit for: Obtain the entropy value of the visual feature map output by the visual sub-model. When the entropy value is lower than the preset threshold, reduce the decision weight of the visual sub-model through the cross-attention module, while increasing the decision weight of the acoustic sub-model. When physical sensor data is detected to be missing, the model is switched to a visual-acoustic dual-modal working mode, and the current wear depth is verified for consistency using the wear depth change trend in historical monitoring data.

[0140] In some embodiments of this disclosure, the apparatus may further include a quantization unit for: Extracting acoustic signature spectral features from acoustic signals in multimodal data; Monitor the energy density changes in the high-frequency bands related to the tire air pumping effect in the acoustic spectrum; When a nonlinear shift in the dominant frequency energy of the acoustic signature spectrum toward the lower frequency band is detected, it is determined that the tire's internal structural health has deteriorated. The audio mutual information entropy of the acoustic signature spectrum is calculated to transform the nonlinear offset characteristics into a wear degree quantification index value, which is then used to determine the hidden wear of the tires of the trackless rubber-tired vehicle in the coal mine.

[0141] In some embodiments of this disclosure, the multimodal fusion deep learning model further includes: The multimodal feature refinement module, which employs the SE channel attention mechanism, is located after the multimodal fusion module. It is used to recalibrate the cross-modal weights of the fused features to increase the weights of modal features that contribute significantly to the current monitoring task.

[0142] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0143] According to the multimodal fusion-based tire health monitoring device for trackless rubber-tired vehicles in coal mines proposed in this disclosure, by acquiring multimodal data of the tire to be monitored and inputting it into a pre-trained multimodal fusion deep learning model, it can output the tire wear category, wear depth value, and heat map indicating the discrimination area, realizing a multi-dimensional joint assessment of wear status. On this basis, the generated heat map is subjected to threshold segmentation and connected component analysis in sequence, and finally the pixel-level wear area corresponding to the wear category in the visual image is determined. Thus, without the need for pixel-level annotation, the image-level classification result is refined into specific wear location information. This not only realizes the full-process automation of tire wear monitoring, effectively avoiding the subjectivity of manual monitoring and the limitations of the underground environment, but also completes wear level judgment, depth estimation, and area positioning simultaneously through an end-to-end deep learning model, significantly improving monitoring efficiency and positioning accuracy.

[0144] Figure 3 This is a block diagram illustrating an apparatus for a method of monitoring the tire health status of a trackless rubber-tired vehicle in a coal mine based on multimodal fusion, according to an exemplary embodiment. For example, apparatus 300 may be an electronic device, such as a mobile phone, computer, digital broadcasting terminal, messaging device, tablet device, personal digital assistant, etc.

[0145] Reference Figure 3 The device 300 may include one or more of the following components: processing component 302, memory 304, power component 306, multimedia component 308, audio component 310, input / output I / O interface 312, sensor component 314, and communication component 316.

[0146] Processing component 302 typically controls the overall operation of device 300, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 302 may include one or more processors 320 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 302 may include one or more modules to facilitate interaction between processing component 302 and other components. For example, processing component 302 may include a multimedia module to facilitate interaction between multimedia component 308 and processing component 302.

[0147] Memory 304 is configured to store various types of data to support the operation of device 300. Examples of such data include instructions for any application or method operating on device 300, contact data, phonebook data, messages, pictures, videos, etc. Memory 304 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0148] The power supply component 306 provides power to the various components of the device 300. The power supply component 306 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device 300.

[0149] Multimedia component 308 includes a screen that provides an output interface between the device 300 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may not only sense the boundaries of touch or swipe actions but also monitor the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 308 includes a front-facing camera and / or a rear-facing camera. When the device 300 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0150] Audio component 310 is configured to output and / or input audio signals. For example, audio component 310 includes a microphone (MIC) configured to receive external audio signals when device 300 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 304 or transmitted via communication component 316. In some embodiments, audio component 310 also includes a speaker for outputting audio signals.

[0151] I / O interface 312 provides an interface between processing component 302 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, start buttons, and lock buttons.

[0152] Sensor assembly 314 includes one or more sensors for providing status assessments of various aspects of device 300. For example, sensor assembly 314 may monitor the on / off state of device 300, the relative positioning of components such as the display and keypad of device 300, changes in the position of device 300 or a component of device 300, the presence or absence of user contact with device 300, the orientation or acceleration / deceleration of device 300, and temperature changes of device 300. Sensor assembly 314 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 314 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 314 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.

[0153] Communication component 316 is configured to facilitate wired or wireless communication between device 300 and other devices. Device 300 can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 316 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 316 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0154] In an exemplary embodiment, the apparatus 300 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0155] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 304 including instructions, which can be executed by a processor 320 of the device 300 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0156] In an exemplary embodiment, a computer program product is also provided, including a computer program that implements the above-described method when executed by the processor 320 of the device 300.

[0157] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the appended claims.

[0158] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A coal mine trackless rubber-tyred vehicle tire health state monitoring method based on multi-modal fusion, characterized in that, include: Acquire multimodal data of the tires of the trackless rubber-tired vehicle to be monitored in the coal mine; the multimodal data includes at least visual images; The multimodal data is input into a pre-trained multimodal fusion deep learning model to obtain the wear category and wear depth value of the tire, as well as a heat map corresponding to the wear category; wherein, the heat map is used to indicate the image region in the model used to distinguish the wear category; The heatmap is subjected to threshold segmentation to obtain a binary image; Perform connected component analysis on the binary image, and determine the pixel-level wear region in the visual image corresponding to the wear category based on the results of the connected component analysis.

2. The coal mine trackless rubber-tyred vehicle tire health state monitoring method based on multi-modal fusion according to claim 1, characterized in that, The multimodal data further includes acoustic signals and physical sensing data; the method further includes: The visual image is subjected to denoising, image enhancement, and tire region cropping to obtain a standardized target tire image; The acoustic signal is denoised, and the denoised acoustic signal is converted into a Mel spectrogram. Time-frequency analysis is performed on the vibration data and / or acceleration data in the physical sensing data to obtain frequency domain features; the tire pressure data and / or tire temperature data in the physical sensing data are normalized to obtain normalized scalar features; the frequency domain features and the scalar features are combined to obtain the target features; The target tire image, Mel spectrogram, and target features are synchronized and aligned in time to obtain aligned multimodal data; The aligned multimodal data is subjected to feature fusion to obtain feature-fused multimodal data; The step of inputting the multimodal data into a pre-trained multimodal fusion deep learning model to obtain the wear category and wear depth value of the tire, as well as a heat map corresponding to the wear category, includes: The multimodal data after feature fusion is input into a pre-trained multimodal fusion deep learning model to obtain the wear category and wear depth value of the tire, as well as the heat map corresponding to the wear category.

3. The coal mine trackless rubber-tyred vehicle tire health state monitoring method based on multi-modal fusion according to claim 2, characterized in that, The multimodal fusion deep learning model includes: The visual sub-model employs a Swing Transformer backbone network to extract multi-scale features of visual images through the self-attention mechanism of the moving window. The visual sub-model includes a channel attention module for weighting different feature channels to enhance the sensitivity to tire wear texture features, and a spatial attention module for focusing on key spatial locations of wear to enhance the localization accuracy of tire wear texture features. Acoustic sub-models are used to extract acoustic features from Mel spectrograms; The physical sub-model is used to extract physical features from physical sensing data. A multimodal fusion module is used to fuse the visual features, acoustic features, and physical features to obtain fused features; A classification output head is used to output the wear category of the tire based on the fused features; A regression output head is used to output the wear depth value of the tire based on the fused features; The weakly supervised localization output head is used to generate a heatmap corresponding to the wear category based on the category activation map or its variant, so as to complete the pixel-level wear area localization based on image-level labels.

4. The method for monitoring the tire health status of trackless rubber-tired vehicles in coal mines based on multimodal fusion according to claim 1, characterized in that, The heat map is obtained using the following steps: Extract the feature map of the preset target convolutional layer in the multimodal fusion deep learning model; The heat map corresponding to the wear category is calculated using the following formula: in, Indicates the type of wear The generated Grad-CAM heatmap, The first target convolutional layer in a deep convolutional neural network Each feature map Wear category Relative to feature map Importance weights Wear category The score.

5. The method for monitoring the tire health status of trackless rubber-tired vehicles in coal mines based on multimodal fusion according to claim 1, characterized in that, The threshold segmentation process of the heatmap to obtain a binary image includes: An adaptive threshold segmentation algorithm is used to calculate the segmentation threshold based on the distribution characteristics of pixel values ​​in the heat map; The pixels in the heatmap whose pixel values ​​are greater than or equal to the segmentation threshold are set as the first value, and the pixels whose pixel values ​​are less than the segmentation threshold are set as the second value, thereby generating the binary image.

6. The method for monitoring the tire health status of trackless rubber-tired vehicles in coal mines based on multimodal fusion according to claim 1, characterized in that, The step of performing connected component analysis on the binary image and determining the pixel-level wear region in the visual image corresponding to the wear category based on the result of the connected component analysis includes: Morphological operations are performed on the binary image to remove noise and broken connectivity regions; the morphological operations include erosion and / or dilation. Connectivity analysis is performed on the binary image after morphological operations to identify at least one connected component. For each connected component, a minimum bounding rectangle or irregular contour is generated to obtain the localization result of the pixel-level wear region.

7. The method for monitoring the tire health status of trackless rubber-tired vehicles in coal mines based on multimodal fusion according to claim 1, characterized in that, After determining the pixel-level wear region in the visual image corresponding to the wear category based on the results of the connected component analysis, the method further includes: Geometric and texture features are extracted from the pixel-level wear region; the geometric features include at least one of area, perimeter and aspect ratio, and the texture features include at least one of gray-level co-occurrence matrix features and local binary pattern features. The geometric and texture features are input into the multimodal fusion deep learning model to calibrate the wear depth value of the pixel-level wear region.

8. The method for monitoring the tire health status of trackless rubber-tired vehicles in coal mines based on multimodal fusion according to claim 1, characterized in that, Before acquiring the multimodal data of the tires of the trackless rubber-tired vehicle to be monitored in the coal mine, the method further includes: Obtain multimodal data samples of tires for trackless rubber-tired vehicles in coal mines, and label each sample with image-level wear category labels and wear depth numerical labels; The multimodal data samples are preprocessed to obtain aligned multimodal training features; An initial multimodal fusion deep learning model is constructed, which includes a visual sub-model, an acoustic sub-model, a physical sub-model, a multimodal fusion module, a classification output head, a regression output head, and a weakly supervised localization output head; The multimodal training features are input into the initial multimodal fusion deep learning model to obtain wear category prediction, wear depth prediction, and heat map prediction. Calculate the classification loss between the image-level wear category label and the wear category predicted value, and calculate the regression loss between the wear depth numerical label and the wear depth predicted value; Calculate the attention consistency loss between the heatmap prediction and the acoustic attention map; the acoustic attention map is an attention map corresponding to the wear category generated by class activation mapping based on the feature map of the acoustic sub-model in the multimodal fusion deep learning model. The model parameters are updated through backpropagation using the weighted sum of the attention consistency loss, classification loss, and regression loss as the total loss, so that the initial multimodal fusion deep learning model can simultaneously learn the wear classification task, the wear depth regression task, and the heatmap generation task.

9. The method for monitoring the tire health status of trackless rubber-tired vehicles in coal mines based on multimodal fusion according to claim 3, characterized in that, The method further includes: The entropy value of the visual feature map output by the visual sub-model is obtained. When the entropy value is lower than a preset threshold, the decision weight of the visual sub-model is reduced by the cross-attention module, while the decision weight of the acoustic sub-model is increased. When the physical sensor data is detected to be missing, the model is switched to a visual-acoustic dual-modal working mode, and the current wear depth is verified for consistency using the wear depth change trend in historical monitoring data.

10. The method according to claim 1, characterized in that, The method also includes: Extract acoustic signature spectral features from the acoustic signals in the multimodal data; Monitor the energy density changes in the high-frequency bands related to the tire air pumping effect in the acoustic spectrum; When a nonlinear shift in the dominant frequency energy of the acoustic signature spectrum toward the lower frequency band is detected, it is determined that the tire has experienced internal structural health degradation. The audio mutual information entropy of the acoustic signature spectrum is calculated to convert the nonlinear offset feature into a wear degree quantification index value, and the wear degree quantification index value is used to determine the hidden wear of the tires of the trackless rubber-tired vehicle in the coal mine.

11. The method for monitoring the tire health status of trackless rubber-tired vehicles in coal mines based on multimodal fusion according to claim 3, characterized in that, The multimodal fusion deep learning model also includes: The multimodal feature refinement module, which employs the SE channel attention mechanism, is located after the multimodal fusion module. It is used to recalibrate the cross-modal weights of the fused features to increase the weights of modal features that contribute significantly to the current monitoring task.