Multi-modal perception data fusion enhancement method and system for coal mine unmanned vehicle
By predicting image quality degradation based on a 3D physical model of the tunnel and predicted microenvironment data, an appropriate visible light and infrared image enhancement strategy was formulated. This solved the problems of insufficient adaptability of multimodal sensing data in coal mines due to changes in the microenvironment and poor cross-modal fusion effect, generating high-quality multimodal fusion sensing results and supporting the safe and stable operation of unmanned vehicles in coal mines.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ANHUI UNIV OF SCI & TECH
- Filing Date
- 2026-01-27
- Publication Date
- 2026-07-21
AI Technical Summary
Existing image enhancement methods are mostly designed for single-modal images or general scenes, lacking dynamic adaptability to the special micro-environment of coal mines. They are difficult to effectively cope with the complex attenuation caused by changes in different micro-environmental parameters to multimodal sensing data. At the same time, in the process of multimodal data fusion, problems such as large differences in features between modes and spatiotemporal asynchrony lead to poor fusion results.
Based on the three-dimensional physical model of the coal mine roadway and the predicted micro-environment data within the preset time zone, a sequence of predicted visible light and infrared image quality degradation indicators is generated, an appropriate image enhancement strategy is formulated, and a multi-modal fusion enhanced perception result is generated through a cross-modal feature fusion plugin, which is then transmitted to the autonomous driving perception and decision center in real time.
The enhancement strategy is forward-looking and targeted, improving the quality of single-modal images in complex downhole environments and reducing the impact of intermodal feature differences and spatiotemporal asynchrony, generating multimodal fusion perception results that more comprehensively reflect downhole environment information.
Smart Images

Figure CN122023144B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image enhancement technology, specifically to a multimodal perception data fusion enhancement method and system for unmanned vehicles in coal mines. Background Technology
[0002] With the deep integration of artificial intelligence and autonomous driving technologies, unmanned vehicles in coal mines, as key equipment for intelligent underground mining, rely on accurate perception of complex tunnel environments for safe and stable operation.
[0003] However, existing image enhancement methods are mostly designed for single-modal images or general scenes, lacking dynamic adaptability to the special micro-environment of coal mines. They are difficult to effectively cope with the complex attenuation caused by changes in different micro-environmental parameters to multimodal sensing data. At the same time, in the process of multimodal data fusion, problems such as large differences in features between modes and spatiotemporal asynchrony lead to poor fusion results. Summary of the Invention
[0004] This application provides a method and system for multimodal perception data fusion and enhancement for unmanned vehicles in coal mines, which solves the technical problems of insufficient adaptability of enhancement strategies and poor cross-modal fusion effect caused by dynamic changes in the microenvironment of existing multimodal perception data in underground coal mines.
[0005] The technical solution to the above-mentioned technical problems in this application is as follows:
[0006] In a first aspect, this application provides a multimodal perception data fusion and enhancement method for unmanned vehicles in coal mines, the method comprising: Based on the three-dimensional physical model of the coal mine roadway and the predicted microenvironment data within the preset time zone, image quality degradation is predicted, generating a sequence of predicted visible light image quality degradation indexes and a sequence of predicted infrared image quality degradation indexes, and formulating strategies for enhancing visible light images and infrared images. When the unmanned vehicle in the coal mine is traveling in the preset time zone, the vehicle perception data is enhanced according to the adapted visible light image enhancement strategy and the adapted infrared image enhancement strategy to generate visible light enhanced image stream and infrared enhanced image stream. The visible light enhanced image stream and infrared enhanced image stream are input into a pre-built cross-modal feature fusion plugin for cross-modal feature alignment, generating a multimodal fusion enhanced perception result, which is then transmitted in real time to the downstream autonomous driving perception and decision-making center.
[0007] Secondly, this application provides a multimodal perception data fusion and enhancement system for unmanned vehicles in coal mines, including: The image processing module is used to predict image quality degradation based on the three-dimensional physical model of the coal mine roadway and the predicted microenvironment data within the preset time zone, generate a sequence of predicted visible light image quality degradation indexes and a sequence of predicted infrared image quality degradation indexes, and formulate an adapted visible light image enhancement strategy and an adapted infrared image enhancement strategy. The image enhancement module is used to enhance the vehicle perception data according to the adapted visible light image enhancement strategy and the adapted infrared image enhancement strategy when the unmanned coal mine vehicle is driving in the preset time zone, and generate visible light enhanced image stream and infrared enhanced image stream. The feature alignment module is used to input the visible light enhanced image stream and the infrared enhanced image stream into a pre-built cross-modal feature fusion plugin for cross-modal feature alignment, generate multimodal fusion enhanced perception results, and transmit them to the downstream autonomous driving perception decision center in real time.
[0008] This application provides one or more technical solutions, which have at least the following technical effects or advantages: This application provides a multimodal perception data fusion and enhancement method and system for unmanned vehicles in coal mines. First, based on a three-dimensional physical model of the coal mine roadway and predicted microenvironment data, image quality attenuation is predicted, generating predicted visible light image quality attenuation index sequences and predicted infrared image quality attenuation index sequences. This allows for the formulation of visible light and infrared image enhancement strategies highly adapted to dynamic microenvironment changes, achieving both forward-looking and targeted enhancement. Second, when the unmanned vehicle is traveling within a preset time zone, real-time image enhancement of the vehicle's perception data is performed according to the formulated adaptive enhancement strategy, generating high-quality visible light enhanced image streams and infrared enhanced image streams, effectively improving the quality of single-modal images in complex underground environments. Finally, the enhanced multimodal image streams are input into a pre-constructed cross-modal feature fusion plugin for feature alignment. Through training and optimization using advanced technologies such as generative adversarial networks, the impact of intermodal feature differences and spatiotemporal asynchrony is reduced.
[0009] Through the above technical solutions, the multimodal fusion enhanced perception results generated in this application can more comprehensively reflect the underground environment information and transmit it to the autonomous driving perception and decision-making center in real time, providing perception assurance for the safe and stable operation of unmanned vehicles in coal mines. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a flowchart illustrating the multimodal perception data fusion and enhancement method for unmanned vehicles in coal mines provided in this application embodiment; Figure 2 This is a schematic diagram of the structure of the multimodal perception data fusion and enhancement system for unmanned vehicles in coal mines provided in the embodiments of this application.
[0012] The components represented by each number in the attached diagram are explained below: Image processing module 11, image enhancement module 12, feature alignment module 13. Detailed Implementation
[0013] This application provides a method and system for multimodal perception data fusion and enhancement for unmanned vehicles in coal mines, which addresses the technical problems of insufficient adaptability of enhancement strategies and poor cross-modal fusion effects caused by dynamic changes in the microenvironment of existing multimodal perception data in underground coal mines.
[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0015] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0016] In the description of this application, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use this application. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that this application can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid unnecessarily obscuring the description of this application. Therefore, this application is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.
[0017] Example 1, as Figure 1As shown in the embodiments of this application, a multimodal perception data fusion and enhancement method for unmanned vehicles in coal mines is provided, including: S10: Based on the three-dimensional physical model of the coal mine roadway and the predicted micro-environment data within the preset time zone, image quality degradation is predicted, a sequence of predicted visible light image quality degradation indexes and a sequence of predicted infrared image quality degradation indexes are generated, and strategies for enhancing visible light images and infrared images are formulated. In this embodiment, the three-dimensional physical model of the tunnel is constructed using laser scanning or tunnel design drawings, and includes spatial feature information such as the tunnel's geometric dimensions, support structure, turning angles, and slope changes. The preset time zone is set according to the coal mine's underground operation plan or environmental change cycle, for example, set to the operation period of the next hour or the next shift.
[0018] The predicted microenvironment data includes predicted dust concentration distribution sequences, predicted water mist concentration distribution sequences, predicted light intensity change sequences, and predicted temperature and humidity change sequences at different locations in coal mine roadways within a preset time zone. These data are obtained by combining historical monitoring data with an environmental prediction model.
[0019] Furthermore, when predicting image quality degradation, for visible light images, the dust concentration, water mist concentration, and light intensity in the predicted microenvironment data are taken as the main influencing factors. Through the constructed visible light image quality degradation prediction model, a sequence of predicted visible light image quality degradation indicators that vary with time and space is generated.
[0020] For infrared images, their quality degradation is mainly affected by the scattering and absorption of infrared radiation by dust and water mist, as well as noise interference caused by changes in ambient temperature. Therefore, based on the predicted dust concentration, water mist concentration and temperature and humidity data, an infrared image degradation prediction model is used to generate a sequence of predicted infrared image quality degradation indicators.
[0021] The attenuation index-visible light enhancement scheme mapping rule base pre-stores the mapping relationship between different attenuation index combinations and corresponding enhancement algorithms and parameters. The attenuation index-infrared enhancement scheme mapping rule base, on the other hand, establishes corresponding rules with attenuation indices based on the characteristics of infrared images. Thus, enhancement schemes and parameters are dynamically selected based on the predicted infrared image quality attenuation index sequence, forming an adapted infrared image enhancement strategy.
[0022] The steps for obtaining predictive microenvironment data include: The coal mine roadways are divided into several roadway sections; The dust concentration sequence of the aforementioned roadway sections within a historical time zone is obtained through monitoring, and a historical dust concentration distribution sequence is generated. The relative humidity and temperature sequences of the aforementioned roadway sections within historical time zones are monitored and obtained. Based on the relative humidity and temperature, the water mist concentration is derived, and a historical water mist concentration distribution sequence is generated. A dust concentration time series predictor and a water mist concentration time series predictor are constructed based on a long short-term memory network. The predicted dust concentration distribution sequence and the predicted water mist concentration distribution sequence are obtained respectively based on the historical dust concentration distribution sequence and the historical water mist concentration distribution sequence, and are used as predicted microenvironment data.
[0023] In this embodiment, firstly, according to the actual structural characteristics of the coal mine roadway, it is divided into several roadway segments with similar environmental characteristics, such as straight roadway segments, curved roadway segments, intersection segments, and slope change segments. The length of each roadway segment is set according to the actual complexity of the roadway. For example, a straight roadway segment can be set to 50 meters per segment, while curved roadway segments and intersection segments are independent roadway segments with complete turning structures or intersection areas.
[0024] Secondly, by deploying dust sensors and temperature and humidity sensors in each roadway section, minute-level monitoring data of dust concentration, temperature, and relative humidity for each roadway section within a historical time zone are collected, forming historical dust concentration sequences, historical temperature sequences, and historical relative humidity sequences. Among them, the distribution of a single historical dust concentration includes the dust concentration and location distribution information of several roadway sections at the same monitoring time.
[0025] Regarding water mist concentration, since water mist in coal mines is mainly generated by air humidity saturation and dust suppression spraying during operations, a water mist concentration derivation model is constructed. This model uses historical relative humidity and temperature sequences as inputs, combined with parameters such as tunnel ventilation volume, to calculate the water mist concentration value at each moment, thereby generating a historical water mist concentration distribution sequence. Specifically, the input parameters of the water mist concentration derivation model include relative humidity, temperature, and tunnel ventilation volume, and the output is the water mist concentration.
[0026] Subsequently, considering the temporal variation characteristics of dust concentration and water mist concentration, a long short-term memory (LSTM) network was used to construct time-series predictors for dust concentration and water mist concentration, respectively. During the model training phase, the historical dust concentration distribution sequence was divided into training and validation sets according to time sequence. The training set data was input into the LSTM model, and hyperparameters such as the number of network layers, the number of hidden units, and the learning rate were adjusted to minimize the mean squared error between the model's predicted values and the actual values on the validation set, thus completing the training of the dust concentration time-series predictor. The same method was used to train the water mist concentration time-series predictor using the historical water mist concentration distribution sequence.
[0027] For example, the following steps are taken to construct a dust concentration time series predictor and a water mist concentration time series predictor based on a long short-term memory network: First, data acquisition involves deploying laser scattering dust sensors in each section of the coal mine roadway to collect historical dust concentration distribution sequences and historical water mist concentration distribution sequences.
[0028] Secondly, regarding model construction, the dust concentration time series predictor adopts a 3-layer LSTM network structure, with an input layer dimension of 1, which is a univariate time series, and the number of neurons in the hidden layer are 64, 32, and 16 respectively. The output layer dimension is 1, the activation function is ReLU, the loss function is mean squared error, the optimizer is Adam, and the learning rate is initially set to 0.001. The water mist concentration time series predictor adopts the same network structure, only the input data is replaced with historical water mist concentration distribution sequences.
[0029] Finally, for model training, the preprocessed historical data is divided into training and validation sets in a 7:3 ratio, with 200 training rounds. After each training round, the validation set loss is calculated. Training stops when the validation set loss no longer decreases for 10 consecutive rounds. The optimal model parameters are saved, and the dust concentration time series predictor and water mist concentration time series predictor are completed.
[0030] Finally, by inputting the time parameters of the preset time zone into the two trained predictors, the predicted dust concentration distribution sequence and the predicted water mist concentration distribution sequence for each tunnel segment at different times within the preset time zone can be obtained.
[0031] Specifically, based on a three-dimensional physical model of a coal mine roadway and predicted microenvironment data within a preset time zone, image quality degradation is predicted, generating a sequence of predicted visible light image quality degradation indices and a sequence of predicted infrared image quality degradation indices, including: Pre-trained visible light image quality degradation prediction plugin and infrared image quality degradation prediction plugin; Based on the three-dimensional physical model of the coal mine roadway, the vehicle driving complexity analysis was performed on the several roadway sections respectively, and several image quality attenuation compensation coefficients were determined according to the complexity analysis results. The predicted microenvironment data is input into the visible light image quality degradation prediction plugin to obtain a first prediction result, and the first prediction result is compensated using the plurality of image quality degradation compensation coefficients to obtain a sequence of predicted visible light image quality degradation indices. The predicted microenvironment data is input into the infrared image quality attenuation prediction plugin to obtain a second prediction result, and the second prediction result is compensated using the plurality of image quality attenuation compensation coefficients to obtain a sequence of predicted infrared image quality attenuation indicators.
[0032] In this embodiment, firstly, a visible light image quality attenuation prediction plugin and an infrared image quality attenuation prediction plugin are pre-trained. Visible light image samples under different micro-environmental conditions in a coal mine, along with corresponding actual image quality attenuation indices, are collected. The image samples and corresponding micro-environmental parameters are used as input, and the actual attenuation indices are used as output. Features are extracted through multiple convolutional and pooling layers, and finally, the predicted visible light image quality attenuation indices are output through a fully connected layer.
[0033] Secondly, based on the 3D physical model of the coal mine roadway, vehicle driving complexity analysis was conducted on each roadway segment, and several image quality attenuation compensation coefficients were determined according to the complexity analysis results. Specifically, the 3D physical model of the roadway includes the geometric features of each roadway segment, such as the length of straight sections, the radius of curvature of curved sections, the number of intersections in crossroads, and the slope of variable slope sections. By quantifying geometric features to evaluate driving complexity, the complexity of each roadway segment was divided into several levels, and a corresponding image quality attenuation compensation coefficient was set for each level. For example, for high-complexity curved and crossroads, due to the rapid changes in camera perspective and the high possibility of environmental occlusion during vehicle movement, the image quality attenuation may be more severe even under the same micro-environment parameters. Therefore, a larger compensation coefficient was set for these sections to enhance and correct the attenuation index in subsequent predictions; while a smaller compensation coefficient was set for low-complexity straight sections.
[0034] Then, the predicted microenvironment data is input into the visible light image quality degradation prediction plugin to obtain the first prediction result. The first prediction result is a preliminary prediction value of the visible light image quality degradation index obtained based on the microenvironment data. Since this preliminary prediction does not consider the additional impact of the driving complexity of the tunnel section itself on image quality degradation, it is necessary to compensate the first prediction result using several image quality degradation compensation coefficients.
[0035] Similarly, the predicted microenvironment data is input into the infrared image quality attenuation prediction plugin to obtain a second prediction result. This result is a preliminary prediction of the infrared image quality attenuation index based on the predicted dust concentration, water mist concentration, and temperature and humidity data. The second prediction result is then compensated using the image quality attenuation compensation coefficient corresponding to each tunnel segment.
[0036] The pre-trained visible light image quality degradation prediction plugin and the infrared image quality degradation prediction plugin include: Based on historical operation monitoring data of coal mine roadways, several sample micro-environment data were collected, and the distribution sequences of historical visible light image quality attenuation index and historical infrared image quality attenuation index under different sample micro-environment data were statistically analyzed to obtain the distribution sequences of visible light image quality attenuation index and infrared image quality attenuation index for several samples. The microenvironmental data of several samples and the distribution sequence of visible light image quality degradation index of several samples are used as the first training data, and K-fold cross-division is performed to obtain K first training sets, where K is an integer greater than or equal to 8; The microenvironmental data of several samples and the distribution sequence of infrared image quality attenuation index of several samples are used as the second training data, and K-fold cross-division is performed to obtain K second training sets; Using the sample microenvironment data as input and the sample visible light image quality degradation index distribution sequence as supervision, the deep learning model is trained to convergence using the K first training sets respectively, generating K visible light degradation prediction units, and integrating them to construct a visible light image quality degradation prediction plugin. Using the sample microenvironment data as input and the sample infrared image quality attenuation index distribution sequence as supervision, the deep learning model is trained to convergence using the K second training sets, generating K infrared attenuation prediction units, and integrating them to construct an infrared image quality attenuation prediction plugin.
[0037] Among them, the visible light image quality degradation index includes at least the average gradient, global contrast, image entropy and color saturation mean, and the infrared image quality degradation index includes at least the thermal contrast, signal-to-noise ratio and temperature gradient variance.
[0038] In this embodiment, firstly, a dataset including different dust concentrations, water mist concentrations, light intensity, temperature, and humidity conditions is selected from historical operational monitoring data of coal mine roadways as sample microenvironment data. Simultaneously, visible light and infrared images are captured underground by unmanned vehicles for the corresponding times in the microenvironment data. For visible light images, the average gradient, global contrast, image entropy, and mean color saturation of each image are calculated using an image quality assessment algorithm and arranged chronologically to form a sample visible light image quality degradation index distribution sequence. For infrared images, considering their thermal imaging characteristics, thermal contrast, signal-to-noise ratio, and temperature gradient variance are calculated, and a sample infrared image quality degradation index distribution sequence is constructed.
[0039] Secondly, the sample microenvironment data is combined with the corresponding sample visible light image quality degradation index distribution sequence to form the first training data. To improve the model's generalization ability, the first training data is subjected to K-fold cross-partitioning, with K set to 8, i.e., divided into 8 subsets. In each partition, 7 subsets are used as the training subset and 1 subset is used as the validation subset, thus obtaining 8 different first training sets. Similarly, the sample microenvironment data is combined with the sample infrared image quality degradation index distribution sequence to form the second training data, and the same K-fold cross-partitioning method is used to obtain 8 second training sets.
[0040] Secondly, during the model training phase, a visible light image quality degradation prediction plugin is constructed. The deep learning model is trained using the micro-environmental data of each first training set as input and the distribution sequence of visible light image quality degradation indicators as supervised output. This deep learning model employs a neural network structure with multiple fully connected layers. The number of neurons in the input layer corresponds to the feature dimension of the sample micro-environmental data. The intermediate layers use activation functions such as ReLU for non-linear mapping, and the number of neurons in the output layer corresponds to the number of visible light image quality degradation indicators; for example, four indicators result in four neurons in the output layer. Each first training set corresponds to one visible light degradation prediction unit. After multiple rounds of iterative training, the model's loss function (e.g., mean squared error) converges on the validation subset. Following this method, eight independent visible light degradation prediction units can be obtained using eight first training sets. Finally, the eight visible light degradation prediction units are integrated, for example, through a weighted average or voting mechanism, to form the final visible light image quality degradation prediction plugin.
[0041] Furthermore, the infrared image quality degradation prediction plugin adopts a similar construction process as the visible light plugin. Using the sample microenvironment data from the second training set as input and the sample infrared image quality degradation index distribution sequence as supervision, eight infrared degradation prediction units are trained. The network structure of the units can be fine-tuned according to the characteristics of the infrared image indices, such as adjusting the number of neurons or activation functions in the intermediate layers. Similarly, the eight trained infrared degradation prediction units are integrated to form the infrared image quality degradation prediction plugin.
[0042] By integrating multiple models, the prediction bias of a single model can be effectively reduced, and the accuracy and robustness of predicting the image quality degradation trend under different micro-environmental conditions can be improved.
[0043] Furthermore, based on the three-dimensional physical model of the coal mine roadway, vehicle driving complexity analysis is performed on the several roadway segments respectively. Based on the complexity analysis results, several image quality attenuation compensation coefficients are determined, including: Based on the three-dimensional physical model of the coal mine roadway, the geometric throughput complexity, motion control complexity and perception and positioning complexity of several roadway segments are evaluated respectively, and the multi-dimensional evaluation results are weighted and fitted to obtain several vehicle driving complexities. Calculate the ratio of the difference between the vehicle driving complexity and the preset standard vehicle driving complexity to the first ratio of the preset standard vehicle driving complexity, and add the first ratio to 1 as the image quality attenuation compensation coefficient. Several image quality attenuation compensation coefficients are calculated based on the several vehicle driving complexities.
[0044] In this embodiment, firstly, for each tunnel segment, geometric parameters for complexity assessment are extracted from the three-dimensional physical model of the tunnel. For geometric throughput complexity assessment, factors such as the tunnel segment's clearance dimensions, minimum turning radius, number of branches at intersections, and their angles are mainly considered.
[0045] For example, when the width of a lane section is less than a preset threshold of 3 meters, or the turning radius of a curve section is less than 1.2 times the minimum turning radius of a vehicle, the geometric passability complexity level increases; for each additional branch in an intersection, the complexity level increases accordingly.
[0046] Secondly, the motion control complexity assessment combines the slope of the tunnel section, the road surface smoothness obtained through historical maintenance records or surface roughness parameters in the 3D model, and the curvature change rate of the curve section. When the absolute value of the slope is greater than 5° or the road surface smoothness deviation exceeds 10cm, the motion control difficulty increases significantly, and the complexity level rises. The perception and positioning complexity assessment focuses on analyzing the distribution of obstructions, lighting conditions, and the number of feature points within the tunnel section, such as wall textures, signs, and other visual features that can be used for positioning. When the proportion of obstructions exceeds 30% or the effective feature point density is less than 2 per square meter, the perception and positioning complexity level increases.
[0047] Secondly, the evaluation results of the geometric passability complexity, motion control complexity, and perception and positioning complexity are quantified and scored. For example, the complexity of each dimension is divided into 1-5 levels, corresponding to scores of 1-5. Based on the driving requirements of unmanned vehicles in coal mines and historical accident statistics, different weights are assigned to the three dimensions: geometric passability complexity is weighted at 0.4, motion control complexity at 0.3, and perception and positioning complexity at 0.3. The comprehensive vehicle driving complexity value for each roadway segment is calculated using the weighted summation formula: "Vehicle driving complexity = Geometric passability complexity score × 0.4 + Motion control complexity score × 0.3 + Perception and positioning complexity score × 0.3".
[0048] Next, a preset standard vehicle driving complexity is set. This standard value is determined based on historical evaluation data of a large number of straight road sections; for example, the standard value is set to 2 points. The difference between the vehicle driving complexity of each road segment and the standard value is calculated. If the difference is positive, it indicates that the complexity of the road segment is higher than the standard; if it is negative, it is lower than the standard. This difference is divided by the preset standard vehicle driving complexity to obtain a first ratio, which is used to characterize the compensation strength. The first ratio is then added to 1 to obtain the image quality attenuation compensation coefficient.
[0049] For example, the vehicle driving complexity score of a certain curve section is 4 points, which is 0.1 points different from the standard value of 3.9 points. The first ratio is 0.1 / 3.9≈0.0256. Therefore, the image quality attenuation compensation coefficient is 1+0.0256≈1.0256, which means that the predicted result of the image quality attenuation index in this curve section needs to be compensated by multiplying the preliminary predicted value by 1.0256.
[0050] Specifically, the predicted microenvironment data is input into the visible light image quality degradation prediction plugin to obtain a first prediction result, and the first prediction result is compensated using the plurality of image quality degradation compensation coefficients to obtain a sequence of predicted visible light image quality degradation indices, including: The prediction complexity coefficient is obtained by evaluating the prediction complexity coefficient based on the predicted dust concentration distribution sequence and the predicted water mist concentration distribution sequence. Q is obtained by multiplying the predicted complexity coefficient by the second ratio of the maximum historical predicted complexity coefficient of the coal mine roadway recorded in the historical time range by K and taking the integer part. In the process of calculating Q, if Q is less than 2, then Q is equal to 2; if Q is greater than K, then Q is equal to K. Q units are randomly selected from the K visible light attenuation prediction units of the visible light image quality attenuation prediction plugin. Visible light image quality attenuation is predicted based on the predicted microenvironment data. The average of the Q prediction results is calculated to obtain the distribution sequence of the predicted visible light image quality attenuation index. The predicted visible light image quality attenuation index distribution sequence is obtained by compensating the predicted visible light image quality attenuation index distribution sequence with the aforementioned image quality attenuation compensation coefficients, and by calculating the mean of the compensated visible light image quality attenuation index distribution at the same time.
[0051] In this embodiment, firstly, a calculation model for the prediction complexity coefficient is constructed based on the predicted dust concentration distribution sequence and the predicted water mist concentration distribution sequence. Specifically, the predicted values of dust concentration and water mist concentration are normalized respectively. For example, the dust concentration value is mapped to the interval [0,1], where 0 represents a dust-free state and 1 represents a severely dusty environment; similarly, the water mist concentration value is also mapped to the interval [0,1].
[0052] Furthermore, a weighted summation method is used to calculate the prediction complexity coefficient. The weights are determined based on historical data analysis of the impact of both on the quality degradation of visible light images. For example, the weight of dust concentration is set to 0.6 and the weight of water mist concentration is set to 0.4. That is, the prediction complexity coefficient = 0.6 × normalized dust concentration + 0.4 × normalized water mist concentration.
[0053] Secondly, the historical maximum prediction complexity coefficient is extracted from the historical operation monitoring data of coal mine roadways. This value is the maximum prediction complexity coefficient calculated over a period of time and is used to relativize the current prediction complexity coefficient. The current prediction complexity coefficient is divided by the historical maximum prediction complexity coefficient to obtain a second ratio, which reflects the relative position of the current microenvironment complexity in the historical data.
[0054] Next, the second ratio is multiplied by K and rounded down, where K is the previously set value of 8, to obtain the Q value. Q represents the number of visible light attenuation prediction units involved in the prediction. To ensure the stability and reliability of the prediction, the range of Q is set to [2, K]. That is, when the calculated Q is less than 2, Q is forced to be equal to 2; when Q is greater than K, Q is forced to be equal to K. For example, if the second ratio is 0.3 and K=8, then Q=0.3×8=2.4, which is rounded down to 2; if the second ratio is 0.9, Q=0.9×8=7.2, which is rounded down to 7, both of which are within the valid range.
[0055] Then, from the eight visible light attenuation prediction units included in the visible light image quality attenuation prediction plugin, Q units are randomly selected. Each selected unit independently predicts visible light image quality attenuation indicators using predicted microenvironment data, including predicted dust concentration, water mist concentration, light intensity, temperature, and humidity, and outputs its own prediction results. Subsequently, the arithmetic mean of the Q prediction results is calculated to obtain a comprehensive distribution sequence of predicted visible light image quality attenuation indicators. This sequence includes the trends of average gradient, global contrast, image entropy, and mean color saturation over time.
[0056] Finally, based on the image quality attenuation compensation coefficients of each roadway segment obtained from the analysis of the three-dimensional physical model of the coal mine roadway, and based on the correspondence of roadway markers, the above-mentioned predicted visible light image quality attenuation index distribution sequence is compensated segment by segment.
[0057] Specifically, during vehicle movement, the current roadway segment is determined in real time. The compensation coefficient corresponding to that segment is then used, and the corresponding index values in the predicted index distribution sequence are multiplied by this compensation coefficient to obtain the compensated visible light image quality degradation index distribution. If the vehicle is in the transition zone of the roadway segment at the same time or the compensation coefficient changes, the average of all applicable compensated index distributions at that time is calculated again to form the final predicted visible light image quality degradation index sequence.
[0058] Furthermore, adaptive visible light image enhancement strategies and adaptive infrared image enhancement strategies are formulated, including: Establish a rule base for mapping attenuation index to visible light enhancement scheme, and formulate a sequence of visible light image enhancement compensation schemes based on the predicted visible light image quality attenuation index sequence as an adaptation visible light image enhancement strategy. A rule base for mapping attenuation index to infrared enhancement scheme is established, and an infrared image enhancement compensation scheme sequence is formulated based on the predicted infrared image quality attenuation index sequence as an adapted infrared image enhancement strategy.
[0059] In this embodiment, firstly, a mapping rule base for attenuation indices and visible light enhancement schemes is constructed based on existing technologies. This rule base, based on extensive historical image enhancement experimental data and expert experience, matches corresponding enhancement algorithms and parameter combinations to different indices and degrees of visible light image quality attenuation. The input dimension consists of a multi-dimensional feature vector composed of 4-5 key attenuation indices, and the output action is a predefined enhancement processing pipeline, including algorithm combinations and parameter presets.
[0060] Secondly, based on the obtained sequence of predicted visible light image quality degradation indices, the average gradient, global contrast, image entropy, and average color saturation are extracted at each time point in the sequence. These index values are then input into a mapping rule base for degradation indices and visible light enhancement schemes. The corresponding visible light image enhancement compensation schemes are retrieved from the rule base using fuzzy or exact matching. Arranging the enhancement schemes corresponding to each time point in chronological order forms a sequence of visible light image enhancement compensation schemes. This sequence represents the adaptive visible light image enhancement strategy, capable of dynamically adjusting the enhancement algorithm and parameters according to changes in the microenvironment.
[0061] Furthermore, for infrared images, a similar method is used to establish a mapping rule base for attenuation index-infrared enhancement scheme. Considering that infrared images mainly reflect the thermal radiation information of objects, their enhancement schemes need to be designed for indices such as thermal contrast, signal-to-noise ratio, and temperature gradient variance. Thermal contrast is defined as the ratio of the difference between the average temperature of the target area and the average temperature of the background area to the background temperature. When the thermal contrast is below 0.15, the rule base recommends using an enhancement algorithm based on morphological top-hat transformation to highlight the details of high-temperature areas through top-hat operations. If the signal-to-noise ratio is low, such as below 10dB, a method combining wavelet thresholding denoising and adaptive gain control is adopted. The wavelet threshold function selects a hard threshold, and the threshold value is calculated based on the specific value of the signal-to-noise ratio using the empirical formula "threshold = 0.5 × background standard deviation × sqrt(2 × ln(number of image pixels))". For cases with excessively high temperature gradient variance, such as greater than 5℃², bilateral filtering is used to smooth the image after enhancement to avoid artifacts. The spatial domain standard deviation and gray-scale domain standard deviation of the filter kernel are adaptively adjusted according to the magnitude of the gradient variance. For every 1℃² increase in gradient variance, the spatial domain standard deviation increases by 0.5 pixel units.
[0062] Similarly, when multiple infrared attenuation indicators deteriorate simultaneously, the rule base will invoke a composite enhancement strategy. For example, when thermal contrast and signal-to-noise ratio are both low, wavelet denoising is performed first, followed by the application of a local contrast enhancement algorithm. Then, based on the predicted infrared image quality attenuation indicator sequence, the corresponding infrared image enhancement compensation scheme for each time moment is retrieved from the rule base to form an adapted infrared image enhancement strategy.
[0063] S20: When the unmanned vehicle in the coal mine is traveling in the preset time zone, the vehicle perception data is enhanced according to the adapted visible light image enhancement strategy and the adapted infrared image enhancement strategy to generate a visible light enhanced image stream and an infrared enhanced image stream. In this embodiment, when the unmanned coal mine vehicle is traveling within a preset time zone, the onboard perception system acquires visible light image streams and infrared image streams in real time. The system's internal timing controller triggers the execution of adapted visible light image enhancement strategies and adapted infrared image enhancement strategies based on the correspondence between the current travel time and the preset time zone. For the visible light image stream, the corresponding enhancement scheme is matched frame-by-frame on the time axis according to the generated visible light image enhancement compensation scheme sequence.
[0064] For example, if the predicted visible light image quality degradation index at a certain moment shows that the average gradient is within the preset normal range of 0.4-0.8 and the global contrast is within the preset normal range of 0.3-0.7, and if the average gradient drops to 0.2 and the global contrast drops to 0.15, then the enhancement scheme corresponding to this index combination in the rule base is called. This scheme includes first performing Retinex enhancement to improve the uneven illumination and increase the global contrast to about 0.4, and then using an unsharpened mask algorithm to increase the average gradient to above 0.5. Specific parameters include setting the Gaussian blur kernel size to 3×3 and the gain coefficient to 1.2.
[0065] For the infrared image stream, frame-by-frame processing is performed according to the infrared image enhancement compensation scheme sequence in the adapted infrared image enhancement strategy. Assuming that the predicted thermal contrast of the infrared image at a certain moment is 0.12, which is lower than the 0.15 threshold and the signal-to-noise ratio is 8dB, which is lower than the 10dB threshold, wavelet thresholding is performed first. The threshold is calculated based on the number of image pixels, such as 640×512=327680 pixels. If the background standard deviation is 20, then the threshold = 0.5×20×sqrt(2×ln(327680))≈0.5×20×sqrt(2×12.7)≈0.5×20×sqrt(25.4)≈0.5×20×5.04≈50.4. The high-frequency coefficients after wavelet decomposition are processed using a hard thresholding function. After denoising, an enhancement algorithm based on morphological top-hat transformation is applied. The structural element is a 5×5 rectangular structure to highlight the details of the high-temperature target area and improve the thermal contrast to above 0.2.
[0066] After the above enhancement processing, optimized visible light enhanced image streams and infrared enhanced image streams are output respectively, ensuring that the visual perception system of unmanned vehicles can acquire clear and effective image data under different micro-environmental conditions.
[0067] S30: Input the visible light enhanced image stream and infrared enhanced image stream into the pre-built cross-modal feature fusion plugin for cross-modal feature alignment, generate multimodal fusion enhanced perception results, and transmit them to the downstream autonomous driving perception decision center in real time.
[0068] In this embodiment, the visible light enhanced image stream and the infrared enhanced image stream are input into a pre-built cross-modal feature fusion plugin for cross-modal feature alignment.
[0069] The aligned bimodal feature maps are fused through an adaptive weighted fusion module, which includes a modality reliability evaluation sub-network. The module quantifies the feature quality of each modality by calculating the entropy and gradient magnitude of the feature maps. The lower the entropy and the higher the gradient magnitude, the richer the effective information contained in the modality feature, and the greater the corresponding weight.
[0070] For example, when the entropy value of a visible light image increases above 0.8 due to dust obstruction, its weight is automatically reduced to 0.3, while the weight of the infrared modality is increased to 0.7. The fused features are aggregated at multiple scales through a pyramid pooling module to generate a fused feature map containing local details and global contextual information. Finally, this map is input to the target detection head and the semantic segmentation head, which output obstacle detection boxes, class probabilities, and alleyway semantic segmentation masks, respectively, which together constitute the multimodal fusion enhanced perception result. This result is transmitted in real time to the autonomous driving perception and decision-making center via in-vehicle Ethernet at a transmission rate of no less than 100Mbps and a latency controlled within 50ms, ensuring that the decision-making center can perform real-time path planning and motion control based on the fused perception result.
[0071] The construction process of the cross-modal feature fusion plugin includes: Several sample visible light enhanced images and several sample infrared enhanced images were collected as input training data, and several standard visible light-infrared image pairs after feature alignment were collected as supervision label data. The generator and discriminator of the generative adversarial network are trained using the input training data and supervised label data respectively until both the generation loss function and the discriminant loss function converge, thus obtaining a cross-modal feature fusion plugin.
[0072] In this embodiment, firstly, visible light and infrared images under different micro-environmental conditions are collected from the actual operation scenario of unmanned vehicles in coal mines as raw sample data. The raw sample data is preprocessed, including image distortion correction, size normalization, contrast stretching, and labeling. The labeling content includes the pixel-level spatial correspondence between the visible light and infrared images, the bounding box and category information of the target object, and the segmentation mask of the semantic region of the roadway. The preprocessed sample data is divided into a training set, a validation set, and a test set, with a ratio of 7:2:1.
[0073] Secondly, a generative adversarial network (GAN) architecture is constructed as the foundational model for the cross-modal feature fusion plugin. The generator adopts an improved U-Net structure, with input consisting of dual-channel feature maps from visible light enhanced images and infrared enhanced images. The encoder extracts multi-scale features through five convolutional layers, followed by batch normalization and the LeakyReLU activation function after each convolution. The decoder uses skip connections to fuse high-resolution features from corresponding layers of the encoder, and finally outputs a feature alignment heatmap of the same size as the input through a 1×1 convolution, which is used to guide the spatial alignment of bimodal features. The discriminator adopts a PatchGAN structure, with input consisting of the aligned feature map output by the generator and the standard aligned feature map from the supervision labels. It uses four convolutional layers to determine whether the input feature map is a true aligned feature pair, and outputs the realism probability for each 16×16 receptive field.
[0074] Next, the generation loss function and the discriminant loss function are designed. The generation loss function consists of three parts: first, L1 loss, which calculates the pixel-level absolute error between the aligned feature map output by the generator and the standard aligned feature map, with a weight of 1.0; second, perceptual loss, which uses a pre-trained ResNet50 network to extract high-level semantic features from the generated feature map and the standard feature map, and calculates the Euclidean distance between the feature vectors, with a weight of 0.5; and third, adversarial loss, which combines the output of the discriminator and uses a minimax strategy to optimize the generator, making the generated aligned feature map as close as possible to the real sample, with a weight of 0.2. The discriminant loss function uses binary cross-entropy loss to optimize the discriminator's ability to distinguish between the real aligned feature map and the generated aligned feature map.
[0075] Finally, model training and optimization were performed. The Adam optimizer was used, with initial learning rates of 0.0002 and 0.0004 for the generator and discriminator, respectively, and β1 set to 0.5. During training, model performance was evaluated on the validation set every 500 iterations, using feature alignment error and object detection accuracy as evaluation metrics. When the validation set loss no longer decreased after 10 consecutive iterations, a learning rate decay strategy was adopted, such as decreasing the learning rate by 0.5 times each time, until the learning rate fell below 1e-6 or the preset maximum number of iterations of 20,000 was reached. After training, the converged generator part was used as the core component of the cross-modal feature fusion plugin, and the model weights were saved.
[0076] In summary, compared to existing technologies, this application achieves accurate prediction of the degradation trend of visible light and infrared image quality by pre-constructing a prediction model of micro-environment data and image quality degradation indicators. Furthermore, it improves the spatiotemporal adaptability of degradation prediction by combining this with a roadway physical model for segment-by-segment compensation. A cross-modal feature fusion plugin is constructed using generative adversarial networks, and adaptive weighted fusion is achieved through modal reliability assessment. Multi-scale feature aggregation further enhances the accuracy of fused perception, effectively solving the problem of single-modal perception being susceptible to interference in the complex environment of coal mines.
[0077] In summary, the embodiments of this application have at least the following technical effects: This application provides a multimodal perception data fusion and enhancement method for unmanned vehicles in coal mines. First, based on a three-dimensional physical model of the coal mine roadway and predicted microenvironment data, image quality attenuation is predicted, generating predicted visible light image quality attenuation index sequences and predicted infrared image quality attenuation index sequences. Then, visible light and infrared image enhancement strategies highly adapted to dynamic microenvironment changes are formulated, achieving both forward-looking and targeted enhancement strategies. Second, when the unmanned vehicle is traveling within a preset time zone, real-time image enhancement of the vehicle's perception data is performed according to the formulated adaptive enhancement strategy, generating high-quality visible light enhanced image streams and infrared enhanced image streams, effectively improving the quality of single-modal images in complex underground environments. Finally, the enhanced multimodal image streams are input into a pre-constructed cross-modal feature fusion plugin for feature alignment. Through training and optimization using advanced technologies such as generative adversarial networks, the impact of intermodal feature differences and spatiotemporal asynchrony is reduced.
[0078] Through the above technical solution, the multimodal fusion enhanced perception results generated by this application can more comprehensively reflect the underground environment information and transmit it to the autonomous driving perception and decision-making center in real time, providing perception guarantee for the safe and stable operation of unmanned vehicles in coal mines, thereby effectively solving the technical problems of insufficient adaptability of enhancement strategies and poor cross-modal fusion effect in the prior art.
[0079] Example 2, as Figure 2 As shown, based on the same inventive concept as the multimodal perception data fusion and enhancement method for unmanned vehicles in coal mines provided in Embodiment 1, this application also provides a multimodal perception data fusion and enhancement system for unmanned vehicles in coal mines, including: Image processing module 11 is used to predict image quality degradation based on the three-dimensional physical model of the coal mine roadway and the predicted micro-environment data within the preset time zone, generate a sequence of predicted visible light image quality degradation index and a sequence of predicted infrared image quality degradation index, and formulate an adapted visible light image enhancement strategy and an adapted infrared image enhancement strategy. Image enhancement module 12 is used to enhance the vehicle perception data according to the adapted visible light image enhancement strategy and the adapted infrared image enhancement strategy when the unmanned vehicle in the coal mine is driving in the preset time zone, and generate visible light enhanced image stream and infrared enhanced image stream. The feature alignment module 13 is used to input the visible light enhanced image stream and the infrared enhanced image stream into the pre-built cross-modal feature fusion plugin for cross-modal feature alignment, generate multimodal fusion enhanced perception results, and transmit them to the downstream autonomous driving perception decision center in real time.
[0080] Furthermore, in one embodiment of the application, the step of obtaining the predicted microenvironment data includes: The coal mine roadways are divided into several roadway sections; The dust concentration sequence of the aforementioned roadway sections within a historical time zone is obtained through monitoring, and a historical dust concentration distribution sequence is generated. The relative humidity and temperature sequences of the aforementioned roadway sections within historical time zones are monitored and obtained. Based on the relative humidity and temperature, the water mist concentration is derived, and a historical water mist concentration distribution sequence is generated. A dust concentration time series predictor and a water mist concentration time series predictor are constructed based on a long short-term memory network. The predicted dust concentration distribution sequence and the predicted water mist concentration distribution sequence are obtained respectively based on the historical dust concentration distribution sequence and the historical water mist concentration distribution sequence, and are used as predicted microenvironment data.
[0081] Furthermore, in one embodiment, image quality degradation prediction is performed based on a three-dimensional physical model of the coal mine roadway and predicted microenvironment data within a preset time zone, generating a predicted visible light image quality degradation index sequence and a predicted infrared image quality degradation index sequence, including: Pre-trained visible light image quality degradation prediction plugin and infrared image quality degradation prediction plugin; Based on the three-dimensional physical model of the coal mine roadway, the vehicle driving complexity analysis was performed on the several roadway sections respectively, and several image quality attenuation compensation coefficients were determined according to the complexity analysis results. The predicted microenvironment data is input into the visible light image quality degradation prediction plugin to obtain a first prediction result, and the first prediction result is compensated using the plurality of image quality degradation compensation coefficients to obtain a sequence of predicted visible light image quality degradation indices. The predicted microenvironment data is input into the infrared image quality attenuation prediction plugin to obtain a second prediction result, and the second prediction result is compensated using the plurality of image quality attenuation compensation coefficients to obtain a sequence of predicted infrared image quality attenuation indicators.
[0082] Furthermore, in one embodiment, the pre-trained visible light image quality degradation prediction plugin and the infrared image quality degradation prediction plugin include: Based on historical operation monitoring data of coal mine roadways, several sample micro-environment data were collected, and the distribution sequences of historical visible light image quality attenuation index and historical infrared image quality attenuation index under different sample micro-environment data were statistically analyzed to obtain the distribution sequences of visible light image quality attenuation index and infrared image quality attenuation index for several samples. The microenvironmental data of several samples and the distribution sequence of visible light image quality degradation index of several samples are used as the first training data, and K-fold cross-division is performed to obtain K first training sets, where K is an integer greater than or equal to 8; The microenvironmental data of several samples and the distribution sequence of infrared image quality attenuation index of several samples are used as the second training data, and K-fold cross-division is performed to obtain K second training sets; Using the sample microenvironment data as input and the sample visible light image quality degradation index distribution sequence as supervision, the deep learning model is trained to convergence using the K first training sets respectively, generating K visible light degradation prediction units, and integrating them to construct a visible light image quality degradation prediction plugin. Using the sample microenvironment data as input and the sample infrared image quality attenuation index distribution sequence as supervision, the deep learning model is trained to convergence using the K second training sets, generating K infrared attenuation prediction units, and integrating them to construct an infrared image quality attenuation prediction plugin.
[0083] Among them, the visible light image quality degradation index includes at least the average gradient, global contrast, image entropy and color saturation mean, and the infrared image quality degradation index includes at least the thermal contrast, signal-to-noise ratio and temperature gradient variance.
[0084] Furthermore, in one embodiment, based on the three-dimensional physical model of the coal mine roadway, vehicle driving complexity analysis is performed on the several roadway segments respectively, and several image quality attenuation compensation coefficients are determined according to the complexity analysis results, including: Based on the three-dimensional physical model of the coal mine roadway, the geometric throughput complexity, motion control complexity and perception and positioning complexity of several roadway segments are evaluated respectively, and the multi-dimensional evaluation results are weighted and fitted to obtain several vehicle driving complexities. Calculate the ratio of the difference between the vehicle driving complexity and the preset standard vehicle driving complexity to the first ratio of the preset standard vehicle driving complexity, and add the first ratio to 1 as the image quality attenuation compensation coefficient. Several image quality attenuation compensation coefficients are calculated based on the several vehicle driving complexities.
[0085] Further, the predicted microenvironment data is input into the visible light image quality degradation prediction plugin to obtain a first prediction result, and the first prediction result is compensated using the plurality of image quality degradation compensation coefficients to obtain a sequence of predicted visible light image quality degradation indices, including: The prediction complexity coefficient is obtained by evaluating the prediction complexity coefficient based on the predicted dust concentration distribution sequence and the predicted water mist concentration distribution sequence. Q is obtained by multiplying the predicted complexity coefficient by the second ratio of the maximum historical predicted complexity coefficient of the coal mine roadway recorded in the historical time range by K and taking the integer part. In the process of calculating Q, if Q is less than 2, then Q is equal to 2; if Q is greater than K, then Q is equal to K. Q units are randomly selected from the K visible light attenuation prediction units of the visible light image quality attenuation prediction plugin. Visible light image quality attenuation is predicted based on the predicted microenvironment data. The average of the Q prediction results is calculated to obtain the distribution sequence of the predicted visible light image quality attenuation index. The predicted visible light image quality attenuation index distribution sequence is obtained by compensating the predicted visible light image quality attenuation index distribution sequence with the aforementioned image quality attenuation compensation coefficients, and by calculating the mean of the compensated visible light image quality attenuation index distribution at the same time.
[0086] Furthermore, in one embodiment of the application, an adaptive visible light image enhancement strategy and an adaptive infrared image enhancement strategy are formulated, including: Establish a rule base for mapping attenuation index to visible light enhancement scheme, and formulate a sequence of visible light image enhancement compensation schemes based on the predicted visible light image quality attenuation index sequence as an adaptation visible light image enhancement strategy. A rule base for mapping attenuation index to infrared enhancement scheme is established, and an infrared image enhancement compensation scheme sequence is formulated based on the predicted infrared image quality attenuation index sequence as an adapted infrared image enhancement strategy.
[0087] Furthermore, the construction process of the cross-modal feature fusion plugin includes: Several sample visible light enhanced images and several sample infrared enhanced images were collected as input training data, and several standard visible light-infrared image pairs after feature alignment were collected as supervision label data. The generator and discriminator of the generative adversarial network are trained using the input training data and supervised label data respectively until both the generation loss function and the discriminant loss function converge, thus obtaining a cross-modal feature fusion plugin.
[0088] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0089] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0090] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.
Claims
1. A multimodal perception data fusion and enhancement method for unmanned vehicles in coal mines, characterized in that, The methods include: Based on a 3D physical model of a coal mine roadway and predicted microenvironment data within a preset time zone, image quality degradation is predicted, generating predicted visible light image quality degradation index sequences and predicted infrared image quality degradation index sequences. Adapted visible light image enhancement strategies and adapted infrared image enhancement strategies are also formulated, including: Based on historical operation monitoring data of coal mine roadways, several sample micro-environment data were collected, and the distribution sequences of historical visible light image quality attenuation index and historical infrared image quality attenuation index under different sample micro-environment data were statistically analyzed to obtain the distribution sequences of visible light image quality attenuation index and infrared image quality attenuation index for several samples. A visible light image quality attenuation prediction plugin is obtained by supervising the training of a deep learning model using several sample micro-environment data and several sample visible light image quality attenuation index distribution sequences as training data. An infrared image quality attenuation prediction plugin is obtained by supervising the training of a deep learning model using several sample micro-environment data and several sample infrared image quality attenuation index distribution sequences as training data. The predicted visible light image quality attenuation index sequences and the predicted infrared image quality attenuation index sequences are respectively predicted by the visible light image quality attenuation prediction plugin and the infrared image quality attenuation prediction plugin. When the unmanned vehicle in the coal mine is traveling in the preset time zone, the vehicle perception data is enhanced according to the adapted visible light image enhancement strategy and the adapted infrared image enhancement strategy to generate visible light enhanced image stream and infrared enhanced image stream. The visible light enhanced image stream and the infrared enhanced image stream are input into a pre-built cross-modal feature fusion plugin for cross-modal feature alignment, generating a multimodal fusion enhanced perception result, which is then transmitted in real time to the downstream autonomous driving perception and decision-making center. The steps for obtaining predictive microenvironment data include: The coal mine roadways are divided into several roadway sections; The dust concentration sequence of the aforementioned roadway sections within a historical time zone is obtained through monitoring, and a historical dust concentration distribution sequence is generated. The relative humidity and temperature sequences of the aforementioned roadway sections within historical time zones are monitored and obtained. Based on the relative humidity and temperature, the water mist concentration is derived, and a historical water mist concentration distribution sequence is generated. A dust concentration time series predictor and a water mist concentration time series predictor are constructed based on a long short-term memory network. The predicted dust concentration distribution sequence and the predicted water mist concentration sequence are obtained respectively based on the historical dust concentration distribution sequence and the historical water mist concentration distribution sequence, and are used as predicted microenvironment data. This includes developing adaptive visible light image enhancement strategies and adaptive infrared image enhancement strategies, including: Establish a rule base for mapping attenuation index to visible light enhancement scheme, and formulate a sequence of visible light image enhancement compensation schemes based on the predicted visible light image quality attenuation index sequence as an adaptation visible light image enhancement strategy. A rule base for mapping attenuation index to infrared enhancement scheme is established, and an infrared image enhancement compensation scheme sequence is formulated based on the predicted infrared image quality attenuation index sequence as an adapted infrared image enhancement strategy.
2. The multimodal perception data fusion and enhancement method for unmanned vehicles in coal mines according to claim 1, characterized in that, Based on a three-dimensional physical model of a coal mine roadway and predicted microenvironment data within a preset time zone, image quality degradation is predicted, generating a sequence of predicted visible light image quality degradation indices and a sequence of predicted infrared image quality degradation indices, including: Pre-trained visible light image quality degradation prediction plugin and infrared image quality degradation prediction plugin; Based on the three-dimensional physical model of the coal mine roadway, the vehicle driving complexity analysis was performed on the several roadway sections respectively, and several image quality attenuation compensation coefficients were determined according to the complexity analysis results. The predicted microenvironment data is input into the visible light image quality degradation prediction plugin to obtain a first prediction result, and the first prediction result is compensated using the plurality of image quality degradation compensation coefficients to obtain a sequence of predicted visible light image quality degradation indices. The predicted microenvironment data is input into the infrared image quality attenuation prediction plugin to obtain a second prediction result, and the second prediction result is compensated using the plurality of image quality attenuation compensation coefficients to obtain a sequence of predicted infrared image quality attenuation indicators.
3. The multimodal perception data fusion and enhancement method for unmanned vehicles in coal mines according to claim 2, characterized in that, Pre-trained visible light image quality degradation prediction plugins and infrared image quality degradation prediction plugins include: The microenvironmental data of several samples and the distribution sequence of visible light image quality degradation index of several samples are used as the first training data, and K-fold cross-division is performed to obtain K first training sets, where K is an integer greater than or equal to 8; The microenvironmental data of several samples and the distribution sequence of infrared image quality attenuation index of several samples are used as the second training data, and K-fold cross-division is performed to obtain K second training sets; Using the sample microenvironment data as input and the sample visible light image quality degradation index distribution sequence as supervision, the deep learning model is trained to convergence using the K first training sets respectively, generating K visible light degradation prediction units, and integrating them to construct a visible light image quality degradation prediction plugin. Using the sample microenvironment data as input and the sample infrared image quality attenuation index distribution sequence as supervision, the deep learning model is trained to convergence using the K second training sets, generating K infrared attenuation prediction units, and integrating them to construct an infrared image quality attenuation prediction plugin.
4. The multimodal perception data fusion and enhancement method for unmanned vehicles in coal mines according to claim 3, characterized in that, Visible light image quality degradation metrics include at least average gradient, global contrast, image entropy, and mean color saturation; infrared image quality degradation metrics include at least thermal contrast, signal-to-noise ratio, and temperature gradient variance.
5. The multimodal perception data fusion and enhancement method for unmanned vehicles in coal mines according to claim 2, characterized in that, Based on the three-dimensional physical model of the coal mine roadway, vehicle driving complexity analysis was performed on several roadway segments. Based on the complexity analysis results, several image quality attenuation compensation coefficients were determined, including: Based on the three-dimensional physical model of the coal mine roadway, the geometric throughput complexity, motion control complexity and perception and positioning complexity of several roadway segments are evaluated respectively, and the multi-dimensional evaluation results are weighted and fitted to obtain several vehicle driving complexities. Calculate the ratio of the difference between the vehicle driving complexity and the preset standard vehicle driving complexity to the first ratio of the preset standard vehicle driving complexity, and add the first ratio to 1 as the image quality attenuation compensation coefficient. Several image quality attenuation compensation coefficients are calculated based on the several vehicle driving complexities.
6. The multimodal perception data fusion and enhancement method for unmanned vehicles in coal mines according to claim 3, characterized in that, The predicted microenvironment data is input into the visible light image quality degradation prediction plugin to obtain a first prediction result, and the first prediction result is compensated using the plurality of image quality degradation compensation coefficients to obtain a sequence of predicted visible light image quality degradation indices, including: The prediction complexity coefficient is obtained by evaluating the prediction complexity coefficient based on the predicted dust concentration distribution sequence and the predicted water mist concentration distribution sequence. Q is obtained by multiplying the predicted complexity coefficient by the second ratio of the maximum historical predicted complexity coefficient of the coal mine roadway recorded in the historical time range by K and taking the integer part. In the process of calculating Q, if Q is less than 2, then Q is equal to 2; if Q is greater than K, then Q is equal to K. Q units are randomly selected from the K visible light attenuation prediction units of the visible light image quality attenuation prediction plugin. Visible light image quality attenuation is predicted based on the predicted microenvironment data. The average of the Q prediction results is calculated to obtain the distribution sequence of the predicted visible light image quality attenuation index. The predicted visible light image quality attenuation index distribution sequence is obtained by compensating the predicted visible light image quality attenuation index distribution sequence with the aforementioned image quality attenuation compensation coefficients, and by calculating the mean of the compensated visible light image quality attenuation index distribution at the same time.
7. The multimodal perception data fusion and enhancement method for unmanned vehicles in coal mines according to claim 1, characterized in that, The construction process of the cross-modal feature fusion plugin includes: Several sample visible light enhanced images and several sample infrared enhanced images were collected as input training data, and several standard visible light-infrared image pairs after feature alignment were collected as supervision label data. The generator and discriminator of the generative adversarial network are trained using the input training data and supervised label data respectively until both the generation loss function and the discriminant loss function converge, thus obtaining a cross-modal feature fusion plugin.
8. A multimodal perception data fusion and enhancement system for unmanned vehicles in coal mines, characterized in that: The method for performing the multimodal perception data fusion and enhancement method for unmanned vehicles in coal mines according to any one of claims 1-7 includes: The image processing module is used to predict image quality degradation based on the three-dimensional physical model of the coal mine roadway and the predicted microenvironment data within the preset time zone, generate a sequence of predicted visible light image quality degradation indexes and a sequence of predicted infrared image quality degradation indexes, and formulate an adapted visible light image enhancement strategy and an adapted infrared image enhancement strategy. The image enhancement module is used to enhance the vehicle perception data according to the adapted visible light image enhancement strategy and the adapted infrared image enhancement strategy when the unmanned coal mine vehicle is driving in the preset time zone, and generate visible light enhanced image stream and infrared enhanced image stream. The feature alignment module is used to input the visible light enhanced image stream and the infrared enhanced image stream into a pre-built cross-modal feature fusion plugin for cross-modal feature alignment, generate multimodal fusion enhanced perception results, and transmit them to the downstream autonomous driving perception decision center in real time.