A large model-based production environment autonomous inspection method
By constructing the TD-MIN inspection model and combining it with the diffusion model and the Transformer module, autonomous inspection of the production environment was achieved. This solved the problems of low inspection efficiency, insufficient data fusion, and insufficient decision reliability in existing technologies, and improved the intelligence and robustness of the inspection system.
Patent Information
- Application Number
- CN202511094782.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-08-06
AI Technical Summary
Existing production environment inspection methods suffer from limited functionality, low efficiency, inability to meet the requirements of real-time intelligent autonomous inspection, insufficient data fusion and scenario reconstruction capabilities, limited model robustness and fault tolerance, insufficient uncertainty quantification and decision reliability, and failure to form an autonomous analysis and decision-making closed loop.
A production environment autonomous inspection method based on a large model is constructed, including building a 3D virtual scene, associating time-series data from multimodal sensors, using the TD-MIN inspection large model for autonomous analysis and decision-making, constructing a model fault tolerance mechanism using a diffusion model and a Transformer module, embedding a hierarchical uncertainty quantification network, and realizing data feature extraction and decision analysis.
It enables unstructured risk reasoning under complex working conditions, improves the intelligence and unmanned operation of inspection work, ensures highly reliable decision-making in complex interference environments, supports global optimization and long-term trend analysis, and ensures the continuous and stable operation of the inspection system.
Smart Images

Figure CN120599714B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of production safety informatization, and particularly relates to a production environment autonomous inspection method based on a large model. BACKGROUND
[0002] With the rapid development of industrial intelligence, safety monitoring and efficient operation and maintenance of production environments (such as coal mines, factories, chemical parks, etc.) have become the core demand to ensure production continuity and personnel safety. Traditional production environment inspection mainly relies on manual inspection or simple data collection by single sensors, which has problems such as low efficiency, limited coverage, and delayed abnormal response, and is difficult to adapt to the monitoring needs of complex and dynamic production scenes.
[0003] In the prior art, some production environments have introduced automated inspection equipment (such as robots, drones) and multi-sensor networks to realize monitoring by collecting multi-modal data such as temperature, humidity, vibration, and gas concentration. However, these technologies still have the following limitations:
[0004] Insufficient data fusion and scene restoration capability: Multi-modal sensor data is often stored and analyzed independently, lacking correlation mapping with the physical space of the production environment, making it difficult to intuitively present the scene state behind the data. For example, it is difficult to locate the fault source of a specific device or area by numerical anomalies alone, resulting in low fault diagnosis efficiency.
[0005] Limited model robustness and fault tolerance: Sensors in production environments are susceptible to interference (such as vibration, dust, electromagnetic noise), leading to data loss, anomalies, or distortion. Existing analysis models rely on complete and high-quality data, and lack fault tolerance for noisy data, making them prone to misjudgment or omission.
[0006] Insufficient uncertainty quantification and decision reliability: Dynamic changes in production environments (such as equipment wear, process adjustment, and environmental parameter fluctuations) can cause data distribution drift, making it difficult for existing models to quantify the uncertainty of analysis results. Decision-making strategies lack dynamic confidence adjustment mechanisms, making it difficult to adapt to precise decision-making needs in complex scenarios.
[0007] No autonomous analysis and decision-making loop formed: Most existing technologies still remain at the data collection and anomaly alarm stage, requiring human intervention for fault diagnosis, strategy formulation, and execution, and cannot realize the full-process autonomous loop from data to decision, restricting the improvement of inspection efficiency.
[0008] Therefore, how to construct an autonomous inspection method that can fuse multi-modal time series data, adapt to dynamic production scenarios, have high fault tolerance and uncertainty quantification capability, and realize full-process automation from three-dimensional scene restoration to intelligent decision-making, has become a technical problem to be solved in the field. SUMMARY
[0009] In view of the above problems in the prior art, the production environment autonomous inspection method based on a large model provided by the present application solves the problems of single inspection function, low inspection efficiency and inability to meet real-time and intelligent autonomous inspection requirements in the prior art.
[0010] To achieve the above-mentioned purposes, the technical solution adopted by the present application is as follows: a production environment autonomous inspection method based on a large model, comprising the following steps:
[0011] S1, a three-dimensional virtual scene corresponding to a production environment is built, and real-time collected multi-modal sensor time series data is associated in the three-dimensional virtual scene;
[0012] S2, a large model framework is constructed and trained to obtain a TD-MIN inspection large model;
[0013] The TD-MIN inspection large model comprises an encoding stage and a decoding stage, the encoding stage and the decoding stage are respectively provided with a first model fault-tolerant mechanism module and a second model fault-tolerant mechanism module, and the decoding stage is further provided with a hierarchical uncertainty quantization network; the first model fault-tolerant mechanism module and the second model fault-tolerant mechanism module realize feature extraction of input data and compensation of abnormal data by introducing an incremental noise injection and denoising strategy of a diffusion model; the hierarchical uncertainty quantization network realizes historical rollback and redundancy check of input data feature extraction by statistically quantifying the uncertainty of input data and combining a dynamic confidence adjustment mechanism;
[0014] S3, the TD-MIN inspection large model is used to autonomously analyze and decide the multi-modal sensor time series data in the three-dimensional virtual scene corresponding to the production environment, and output decision analysis data;
[0015] S4, according to the decision analysis data, a decision scene and its corresponding decision strategy are determined and pushed to a corresponding target terminal to realize autonomous inspection of the production environment.
[0016] Further, in the TD-MIN inspection large model in step S2,
[0017] The first model fault-tolerant mechanism module and the second model fault-tolerant mechanism module have the same structure and both comprise an input embedding layer, the output end of the input embedding layer is connected to a plurality of parallel linear layers, the output end of each linear layer is sequentially connected to a fault-tolerant encoding layer and an attention mechanism layer, the output end of each attention mechanism layer is connected to the input end of a fault-tolerant feedforward neural network, and the fault-tolerant feedforward neural network outputs nonlinear transformation data;
[0018] The hierarchical uncertainty quantification network comprises a first UNet module, a first feature extraction layer, an uncertainty quantification layer, a skip connection layer, a second feature extraction layer, a second UNet module and a redundancy check layer connected in sequence, and the redundancy check layer outputs decision analysis data after redundancy check.
[0019] Further, in the step S3, the process of autonomous analysis and decision making by the TD-MIN large model is as follows:
[0020] S301, input the sensor group data as model input data, and perform position, sensor type and point cloud data addressing on the model input data through the embedding layer to obtain encoded sensor group data;
[0021] The sensor group data includes multi-modal sensor time series data and point cloud data thereof in a three-dimensional virtual scene.
[0022] S302, in the encoding stage, after processing the encoded sensor group data through the first fully connected layer, the dependency relationship of the encoded sensor group data in its information sequence is established through the self-attention mechanism layer, and the dynamic confidence adjustment mechanism is introduced to avoid sequence recursion of the network structure in the encoding stage;
[0023] S303, the encoded sensor group data through the self-attention mechanism layer is spliced through the splicing layer, and then input into the first model fault tolerance mechanism module after gradient descent calculation through the second fully connected layer and the third fully connected layer in sequence;
[0024] S304, in the first model fault tolerance mechanism module, the input data is processed and the first diffusion model is gradually injected with noise and denoised, and nonlinear change data is output;
[0025] S305, the nonlinear change data is processed through the first residual network layer, the first standard normalization layer, the first maximum pooling layer, the fourth fully connected layer and the first Softmax layer in sequence to obtain encoding stage output data;
[0026] The encoding stage output data includes compensation abnormal data and feature extraction data.
[0027] S306, in the decoding stage, after processing the encoding stage output data through multiple convolution operations and feedforward neural networks in sequence, the data is input into the second model fault tolerance mechanism module for processing and the second diffusion model is gradually injected with noise and denoised, and nonlinear change data is output;
[0028] S307, the nonlinear change data output by the second model fault tolerance mechanism module is processed through the second residual network layer and the second standard normalization layer in sequence, and then input into the hierarchical uncertainty quantification network;
[0029] S308, in the hierarchical uncertainty quantification network, the input data is modeled with multi-granularity uncertainty to capture the total uncertainty of the input data at different levels, and a redundancy check is performed according to the total uncertainty, and decision analysis data is outputted;
[0030] S309, the decision analysis data is processed by a third residual network layer, a second standard normalization layer, a second maximum pooling layer, a fifth full connection layer, a second Softmax layer and a multi-layer perception MLP in sequence, and the decision analysis data is outputted.
[0031] Further, in the step S302, the implementation process of the dynamic confidence adjustment mechanism is:
[0032] S302-1, an attention weight correction formula is constructed:
[0033]
[0034] In the formula, represents the output result of the confidence-corrected attention mechanism, respectively represent the query , key and value matrices of the sensor group data, represents the dimension of the key vector, represents the confidence mask vector, and the superscript represents the transpose operation, represents the element-wise multiplication;
[0035] S302-2, calculating the self-adaptive adjustment decision threshold
[0036]
[0037] In the formula, represents the initial decision threshold, represents the decay coefficient, represents the total uncertainty at time , and represents the state output at time
[0038] S302-3, judging whether it is true or not;
[0039] If yes, the history rollback mechanism is triggered, the historical data is used to correct the prediction value, and the final prediction value is calculated.
[0040]
[0041] If not, the current prediction value is directly used ;
[0042] wherein, represents the final prediction value at time t, represents a mixing coefficient that determines the weight of the current prediction value and the historical data, represents a length of the historical window, represents a high-confidence prediction value at time t-h
[0043] Further, in the first / second model fault-tolerant mechanism module, the process of progressive noise injection and denoising of the diffusion model is as follows:
[0044] A1. Inject Gaussian noise step by step through Markov chain in the diffusion process of the diffusion model for the original sensor group data;
[0045] wherein, the process of injecting Gaussian noise is represented as:
[0046]
[0047] In the formula, represents the conditional probability distribution of the data at step t when the data at step t-1 is, represents the data at step t generated by Gaussian noise, represents the data at step t generated by Gaussian noise, represents a noise scheduling parameter, represents a Gaussian distribution, represents a covariance; A2. Construct a noise prediction network for reverse process learning to denoise the injected Gaussian noise to approximate the real noise; wherein, the denoising objective function of the noise prediction network is:
[0048]
[0049] wherein, the denoising objective function of the noise prediction network is:
[0050]
[0051] In the formula, represents the actual injected Gaussian noise, represents the noise prediction network, represents the expectation of the time step t , the original data x and the Gaussian noise n original input data representing clean data samples without noise;
[0052] A3、based on the denoising training process, when the sensor group data exists drift or missing, the feature is recovered through the conditional inverse process;
[0053] wherein the formula for recovering the feature through the conditional inverse process is:
[0054]
[0055] wherein, , , , denotes the denoised data at time step , denotes the noise attenuation coefficient at time step , denotes the cumulative attenuation coefficient, denotes the noise attenuation coefficient at time step , denotes the time step index, denotes the standard Gaussian noise vector.
[0056] A3、based on the denoising training process, when the sensor group data exists drift or missing, the feature is recovered through the conditional inverse process;
[0057] wherein the formula for recovering the feature through the conditional inverse process is:
[0058]
[0059] wherein, , , , denotes the denoised data at time step , denotes the noise attenuation coefficient at time step , denotes the cumulative attenuation coefficient, denotes the noise attenuation coefficient at time step , denotes the time step index, denotes the standard Gaussian noise vector.
[0060] In the step S308, the total uncertainty of the input data at different levels is;
[0061]
[0062]
[0063]
[0064]
[0065] wherein, denotes the expected value of the aleatoric uncertainty resulting from the heteroscedastic regression modeling of the input data denotes the mean prediction value outputted by the heteroscedastic regression modeling of the input data denotes the variance prediction value outputted by the heteroscedastic regression modeling of the input data denotes a Gaussian distribution, denotes the observation of the data state of the sensor device group in the production environment;
[0066] denotes the cognitive uncertainty resulting from the Monte Carlo sampling evaluation of the input data denotes the number of Monte Carlo sampling times, denotes the random weight matrix of the -th sampling, denotes the model output under the input data and the random weight denotes the expected value of the output.
[0067] Further, in the step S308, the outputted decision analysis data is:
[0068]
[0069] wherein, denotes an indicator function, denotes the total uncertainty, denotes the adaptive adjustment of the decision threshold in the dynamic confidence adjustment, denotes the initial prediction confidence output of the input denotes a redundancy check function.
[0070] Further, the training loss function of the TD-MIN inspection large model is
[0071]
[0072] wherein, denotes the task loss, denotes the denoising objective function of the noise prediction network in the diffusion model, denotes the total uncertainty, represents a desired function, respectively represent and a trade-off coefficient.
[0073] Further, in the step S4, the decision-making scenario includes:
[0074]
[0075] wherein, represents a set of abnormal, alarm and fault events in a production environment, ; wherein, the priority of the decision-making scenario is sequentially reduced;
[0076] The decision-making strategy corresponding to the decision-making scenario is:
[0077]
[0078] The target terminal pushing rule corresponding to the decision-making scenario is:
[0079]
[0080] wherein, represents a centralized terminal in a production environment, represents a dispatch terminal outside the production environment, represents a mobile terminal APP;
[0081] The dynamic updating equation of the decision-making strategy is:
[0082]
[0083] satisfies:
[0084] wherein, represents a state of the system at time, represents that the real-time pushing delay approaches zero, represents a set of all states of devices in a production environment, represents a set of parameters of the production environment, represents an event-driven decision-making execution function, represents hierarchical logic.
[0085] The present application has the following beneficial effects:
[0086] (1) The TD-MIN inspection big model is obtained by using multi-modal data and fusing a large model framework to construct and train, to analyze and make decisions and drive autonomous inspection reasoning, real-time multi-modal data is directly used to construct a three-dimensional virtual scene environment in the whole production environment, unified analysis, management and inspection of production equipment in the virtual scene, compared with the existing inspection method, not completely dependent on scheduling, analysis or decision-making and other separate systems, breaking through the limitations of traditional threshold alarm, through TD-MIN inspection big model joint modeling, realizing unstructured risk reasoning under complex working conditions, greatly improving the intelligentization and unmannedization of inspection work in different production environments.
[0087] (2) The TD-MIN inspection big model framework proposed in the present application is provided with a model fault tolerance mechanism in the encoding stage and the decoding stage, which constructs a double protection system with noise robustness and dynamic reliability evaluation by fusing diffusion model and Transformer module, adopts the progressive noise injection and denoising training strategy of diffusion model, so that the model can still extract features and compensate data anomalies when the sensor signal drifts or is missing; the hierarchical uncertainty quantization module is embedded in the Transformer encoder module in the decoding stage, the model cognitive uncertainty and data accidental uncertainty are evaluated synchronously through Monte Carlo sampling and heteroscedastic regression, and the attention weight and decision threshold are dynamically adjusted based on real-time confidence, triggering fault tolerance mechanisms such as historical rollback and redundancy check, finally ensuring that the big model framework realizes high reliable decision-making in complex interference environment, while ensuring the accuracy of the data analysis module and the autonomous inspection module of the inspection system, improving the robustness.
[0088] (3) The present application adopts multi-modal data to complete the data analysis, decision-making and inspection function of the inspection process, and cooperates with edge and cloud computing, selects double-channel communication to cooperate with high-performance algorithm servers to deploy lightweight big models, and executes local real-time decision-making; in addition, this system does not rely on a certain sensor or a certain sensor to complete the analysis, decision-making and inspection function, supports global optimization and long-term trend analysis, and ensures the continuous, normal and stable operation of the inspection system in extreme production environment. BRIEF DESCRIPTION OF DRAWINGS
[0089] Figure 1 The production environment autonomous inspection method based on a big model provided by the present application is provided.
[0090] Figure 2 The autonomous inspection system architecture of the coal mine provided by the present application is provided.
[0091] Figure 3 The structure schematic diagram of the TD-MIN inspection big model provided by the present application is provided.
[0092] Figure 4A first / second model fault-tolerant mechanism module structure schematic diagram provided by the present application.
[0093] Figure 5 A hierarchical uncertainty quantification network structure schematic diagram provided by the present application. DETAILED DESCRIPTION
[0094] The specific embodiments of the present application are described below to facilitate those skilled in the art to understand the present application, but it should be clear that the present application is not limited to the scope of the specific embodiments, and for those skilled in the art, it is obvious that various changes are within the spirit and scope of the present application defined and determined by the appended claims, and all inventions utilizing the concept of the present application are within the scope of protection.
[0095] The embodiment of the present application provides a production environment autonomous inspection method based on a large model, as shown in the figure, comprising the following steps: Figure 1
[0096] S1, a three-dimensional virtual scene corresponding to the production environment is built, and real-time collected multi-modal sensor time series data is associated in the three-dimensional virtual scene;
[0097] S2, a large model framework is constructed and trained to obtain a TD-MIN inspection large model;
[0098] The TD-MIN inspection large model includes an encoding stage and a decoding stage, and the encoding stage and the decoding stage are respectively provided with a first model fault-tolerant mechanism module and a second model fault-tolerant mechanism module, and the decoding stage is further provided with a hierarchical uncertainty quantification network; the first model fault-tolerant mechanism module and the second model fault-tolerant mechanism module realize feature extraction of input data and compensation of abnormal data by introducing progressive noise injection and denoising strategy of diffusion model; the hierarchical uncertainty quantification network realizes historical rollback and redundancy check of input data feature extraction by statistically quantifying the uncertainty of input data and combining a dynamic confidence adjustment mechanism;
[0099] S3, using the TD-MIN inspection large model to autonomously analyze and decide the multi-modal sensor time series data in the three-dimensional virtual scene corresponding to the production environment, and outputting decision analysis data;
[0100] S4, according to the decision analysis data, determining a decision scene and a corresponding decision strategy, and pushing to a corresponding target terminal to realize autonomous inspection of the production environment.
[0101] In one specific embodiment of the present application, the production environment is taken as an example of coal mine underground for autonomous inspection:
[0102] To realize the autonomous inspection in the coal mine, first, the inspection system is built, and the specific hardware devices include a 3D vision sensor (time of flight, ToF), an Ethernet controller of a hydraulic support, a GNSS integrated navigation system, a high-performance computing server, a laser radar, a laser scanner, a wireless signal receiver and transmitter, a wireless smoke alarm sensor, a wireless pressure sensor, a wireless inclination sensor, a wireless infrared sensor, a wireless travel sensor, a wireless height sensor, a wireless temperature sensor and a wireless flow sensor.
[0103] The configuration and installation mode of the above hardware devices are as shown in Figure 2 Each hydraulic support is equipped with an Ethernet controller, a wireless sensor cluster (including 2 pressure sensors, 4 wireless inclination sensors, 1 wireless height sensor and 1 wireless travel sensor); one ToF-3D vision camera is installed on the top of each hydraulic support; one coal safety, explosion-proof housing and special protective housing are installed above the coal mining machine, and the housing is internally provided with a high-performance computing server, a GNSS integrated navigation system, a laser radar, a laser scanner, a wireless signal receiver and transmitter, a power supply and other devices, and the laser scanner is installed in the specific protective housing; the support controllers are connected in series through Ethernet, the support controllers can perform bidirectional communication with the high-performance computing server through the wireless signal receiver and transmitter, the high-performance computing server can control the support through the support controller, and the support controller can transmit all support information to the high-performance computing server. Meanwhile, the high-performance computing server can call various sensors, ToF-3D vision cameras and the like according to requirements. According to the requirements of the mine party, the high-performance computing server can transmit the inspection, data analysis, decision and other related information of the fully-mechanized working face to different places such as the underground centralized control, the surface dispatching center and the mobile terminal APP through Ethernet or optical fiber.
[0104] In the inspection system architecture taking the coal mine as an example, S1 in the embodiment includes the following steps:
[0105] S11, the support arrangement and the arrangement of the scraper conveyor of the fully-mechanized working face are arranged parallel to the coal wall, the GNSS integrated navigation system is powered on, the longitude, latitude and height of the current position of the coal mining machine are calculated and the alignment operation is performed;
[0106] S12, the arrangement direction of the scraper conveyor is taken as the X-axis direction, the direction perpendicular to the arrangement of the scraper conveyor is taken as the Y-axis direction, and the direction perpendicular to the top screen of the coal mining machine is taken as the Z-axis direction, a reference coordinate system is established and serves as the coordinate system of the virtual three-dimensional scene of the fully-mechanized working face:
[0107] S13, on the basis of the reference coordinate system, the surrounding environment is scanned by the laser radar and the laser scanner respectively, and three-dimensional environment data is collected;
[0108] The ToF-3D vision camera in the sensor group in the fully mechanized working face collects the RGBD three-dimensional structural environment data around each support;
[0109] In the embodiment, the laser radar and the laser scanner respectively scan the surrounding environment to establish corresponding three-dimensional environment data as , The ToF-3D vision camera collects the RGBD three-dimensional structural environment data around each support as ToF- ;
[0110] S14, in the establishment of the reference coordinate system, according to the three-dimensional environment data and the RGBD three-dimensional structural environment data, the three-dimensional virtual scene of the fully mechanized working face is built;
[0111] In the embodiment, the process of building the corresponding three-dimensional virtual scene of the fully mechanized working face is specifically:
[0112] The RGBD three-dimensional structural environment data, the laser radar data L liDAR The laser scanner data G scanner The Kalman filter algorithm is used for time alignment and synchronization of multi-source data;
[0113] The iterative closest point algorithm is used to register and spatially align the RGBD three-dimensional structural environment data, the radar point cloud information and the scanner point cloud information. The three-dimensional point cloud scene after time alignment and spatial alignment is processed by the method of region growing and point cloud density fusion to realize fine noise layering and denoising. The end-to-end real-time point cloud scene segmentation processing is completed through the hierarchical feature extraction and local region aggregation in PointNet++. The segmented and standardized point cloud scene is fused and reconstructed through the elastic network regression algorithm, and finally the three-dimensional virtual scene of the entire underground fully mechanized working face under the reference coordinate system and the reference coordinate origin is generated.
[0114] S15, whether the built three-dimensional virtual scene is consistent with the real scene of the fully mechanized working face is judged;
[0115] If yes, go to step S16;
[0116] If not, return to step S13;
[0117] S16, the multi-modal time sequence sensor data collected by the sensor group in the fully mechanized working face in real time is displayed at the positions of the sensors in the three-dimensional virtual scene;
[0118] In the embodiment, the multi-modal time sequence sensor data collected by the sensor group in real time includes the corresponding time sequence sensor data collected by the pressure, inclination, infrared and travel sensors, which are respectively pressure , stroke , tilt angle , wherein n represents the number of hydraulic supports.
[0119] In step S2 of the embodiment of the present application, a Transform-Diffusion Integrated Mine Inspection Model large model network structure suitable for autonomous inspection analysis and decision of the present application is built and distilled based on multi-modal data and fusion of a self-attention mechanism module of a Transformers framework and a UNet module in Diffusion in the large model framework, and the Transform-Diffusion Integrated Mine Inspection Model large model network structure is trained by using all possible data information of the sensors prepared in advance and parameter data of the system equipment when stopping, running and failing, to obtain a TD-MIN inspection large model as shown in Figure 3 .
[0120] In the TD-MIN inspection large model in the embodiment, the first model fault-tolerant mechanism module and the second model fault-tolerant mechanism module are respectively arranged in the encoding stage and the decoding stage, which constructs a double protection system with noise robustness and dynamic credibility evaluation by fusing the diffusion model and the self-attention mechanism module in the Transformers framework. On the one hand, the progressive noise injection and denoising training strategy of the diffusion model is adopted to enable the model to still extract features and compensate data anomalies when the sensor signal drifts or is missing; the decoding stage is also provided with an embedded hierarchical uncertainty quantification network built by combining the UNet module in the Diffusion framework to trigger fault-tolerant mechanisms such as history rollback and redundancy check.
[0121] In the embodiment, as shown in Figure 4 , the first model fault-tolerant mechanism module and the second model fault-tolerant mechanism module are the same in structure, and each includes an input embedding layer, the output end of the input embedding layer is connected to a plurality of parallel linear layers, the output end of each linear layer is connected to a fault-tolerant encoding layer and an attention mechanism layer in sequence, the output end of each attention mechanism layer is connected to the input end of a fault-tolerant feedforward neural network, and the fault-tolerant feedforward neural network outputs nonlinearly transformed data.
[0122] In the embodiment, as shown in Figure 5As shown, the hierarchical uncertainty quantification network comprises a first UNet module, a first feature extraction layer, an uncertainty quantification layer, a skip connection layer, a first feature extraction layer, a second UNet module and a redundancy check layer connected in sequence, and the redundancy check layer outputs decision analysis data after redundancy check; Wherein, the first feature extraction layer comprises Conv1, Conv2, Conv3, ReLU1, ReLU2 and two MaxPooling connected in sequence, and the second feature extraction layer comprises Upsample1, Conv1, Upsample2, Conv2, Upsample3, Conv3 and two MaxPooling connected in sequence.
[0123] In the embodiment of the application, based on Figure 3 Based on the model architecture shown, in step S3 of the embodiment, the process of autonomous analysis and decision making by the TD-MIN large model is as follows:
[0124] S301, the sensor group data is used as model input data, and the model input data is respectively addressed by position, sensor type and point cloud data through the embedding layer to obtain encoded sensor group data;
[0125] The sensor group data includes multi-modal sensor time series data and point cloud data in a three-dimensional virtual scene;
[0126] S302, in the encoding stage, after the encoded sensor group data is processed by the first fully connected layer, the internal dependency of the encoded sensor group data in its information sequence is established through the self-attention mechanism layer, and the dynamic confidence adjustment mechanism is introduced to avoid sequence recursion of the network structure in the encoding stage;
[0127] Specifically, in the processing process of the self-attention mechanism layer, according to the established internal dependency and the correlation score of the current data and other types of sensor data, the weight is dynamically allocated according to the weighted aggregation of the context information according to the correlation score, and the matrix calculation is performed on the encoded sensor group data according to the allocated weight, and the position relationship of all sensor group data is processed;
[0128] S303, the encoded sensor group data through the self-attention mechanism layer is spliced through the splicing layer, and after the gradient descent calculation of the second fully connected layer and the third fully connected layer in sequence, it is input into the first model fault tolerance mechanism module;
[0129] S304, in the first model fault tolerance mechanism module, the input data is processed and the progressive noise injection and denoising of the first diffusion model are carried out, and the nonlinear change data is output;
[0130] S305, sequentially through the first residual network layer, the first standard normalization layer, the first maximum pooling layer, the fourth full connection layer and the first Softmax layer, the nonlinear change data is processed to obtain the encoding stage output data;
[0131] Wherein, the gradient vanishing and explosion is relieved by the skip connection in the first residual network, the original information of the input data is retained to improve the feature reuse, and the model convergence is accelerated;
[0132] The encoding stage output data includes compensation abnormal data and feature extraction data;
[0133] S306, in the decoding stage, the encoding stage output data is sequentially processed through multiple convolution operations (Conv1~Conv4) and feedforward neural network, and then input to the second model fault tolerance mechanism module for processing and the second time diffusion model progressive noise injection and denoising, output nonlinear change data;
[0134] S307, the nonlinear change data output by the second model fault tolerance mechanism module is sequentially processed through the second residual network layer and the second standard normalization layer, and then input to the hierarchical uncertainty quantization network;
[0135] Wherein, the same as the first model fault tolerance mechanism module in the encoding stage, the gradient vanishing and explosion is relieved by the skip connection in the second residual network, the original information of the input data is retained to improve the feature reuse, and the model convergence is accelerated;
[0136] S308, in the hierarchical uncertainty quantization network, the input data is modeled with multiple granularity uncertainty to capture the total uncertainty of the input data at different levels, and according to the redundancy check, the decision analysis data is output;
[0137] Specifically, in the hierarchical uncertainty quantization network, the statistical uncertainty and cognitive uncertainty of different levels in the feature information are captured by the UNet module, multiple convolution operations, uncertainty quantization layer, skip connection layer and redundancy check layer for multiple granularity uncertainty modeling;
[0138] S309, sequentially through the third residual network layer, the second standard normalization layer, the second maximum pooling layer, the fifth full connection layer, the second Softmax layer and the multilayer perception MLP, the decision analysis data is processed to output the decision analysis data;
[0139] Specifically, the feature extraction data of the addition history rollback and redundancy check mechanism are connected to the third residual network, and then normalized and standardized. Finally, the corresponding data and corresponding data features are output by the second Softmax layer and the multi-layer perception MLP (MLP) to obtain the corresponding conditions of the normal, different anomalies, alarms and failures of the working face under the current data, as the decision analysis data.
[0140] In step S302 of the embodiment, the implementation process of the dynamic confidence adjustment mechanism is as follows:
[0141] S302-1, constructing an attention weight correction formula:
[0142]
[0143] In the formula, represents the output result of the confidence-corrected attention mechanism, realizing adaptive feature fusion of the sensor group data. The output result is obtained by querying the query and calculating the correlation with the key and then performing weighted aggregation on the value . , respectively, represent the query , key and value matrices of the sensor group data, represents the dimension of the key vector, represents the confidence mask vector, and the superscript represents the transpose operation, represents the element-wise multiplication.
[0144] S302-2, calculating the adaptively adjusted decision threshold .
[0145]
[0146] In the formula, represents the initial decision threshold, which is preset according to the task requirement, represents the decay coefficient, which is used to control the rate of change of the uncertainty of the decision threshold, represents the total uncertainty at time , which integrates the cumulative effect of historical uncertainty, represents the state output at time , which is used to feed back parameter adjustment.
[0147] S302-3, determining whether is true.
[0148] If so, the historical rollback mechanism is triggered, historical data is used to correct the predicted value, and the final predicted value is calculated. ;
[0149]
[0150] If not, then use the current predicted value directly. That is, when the conditions for triggering the historical rollback mechanism are not met, the prediction results at the current moment are fully trusted and no historical data is used for correction.
[0151] in, express The final predicted value at any given time is obtained by using the current predicted value. With history The mixed prediction result is obtained by calculating the weighted average of the predicted values at each time point. This represents the mixing coefficient that determines the weights of the current forecast value and historical data. Indicates the length of the history window. Representing historical moments The high confidence prediction value.
[0152] In this embodiment, in the first / second model fault-tolerant mechanism module, the nonlinearly changing data is output after processing through the input embedding layer, linear layer, fault-tolerant coding layer, attention mechanism layer, and fault-tolerant feedforward neural network. Simultaneously, during this process, the progressive noise injection and denoising process of the diffusion model is completed, ensuring that this data can stably extract features and compensate for abnormal data based on the similarity of other similar sensor data when certain sensor signals drift or are missing. Specifically, the progressive noise injection and denoising process of the diffusion model is as follows:
[0153] A1. Gaussian noise is gradually injected into the original sensor group data through a Markov chain during the diffusion process of the diffusion model.
[0154] The process of injecting Gaussian noise is represented as follows:
[0155]
[0156] In the formula, Indicates the first Step data At that time, the first Step data The conditional probability distribution, Indicates the first The data generated by stepping through Gaussian noise, Indicates the first The data generated by stepping through Gaussian noise, Indicates noise scheduling parameters, Indicates a Gaussian distribution. Represents covariance;
[0157] A2. Construct a noise prediction network based on reverse process learning to denoise the injected Gaussian noise and train it to approximate the real noise.
[0158] Among them, the denoising objective function of the noise prediction network for:
[0159]
[0160] In the formula, This represents the actual injected Gaussian noise. Represents a noise prediction network. Indicates time step Raw data and Gaussian noise Expectations The original input data represents a clean, noise-free data sample.
[0161] A3. Based on the denoising training process, when there is drift or missing data in the sensor group, features are recovered through a conditional inverse process.
[0162] The formula for recovering features through the inverse conditional process is as follows:
[0163]
[0164] In the formula, , , , Indicates time step Denoising data, This represents the noise attenuation coefficient, used for signal recovery during the reverse denoising process. Represents the cumulative decay coefficient, characterizing the decay from the initial value to the... Total noise attenuation of the step Indicates at time step The noise attenuation coefficient, This represents the time step index (a summation variable used for cumulative products). Describes a standard Gaussian noise vector that satisfies , used for random sampling in the reverse process.
[0165] In step S308 of this embodiment, the total uncertainty of the input data at different levels for;
[0166]
[0167]
[0168]
[0169]
[0170] In the formula, Indicates input data Random uncertainty obtained by heteroscedastic regression modeling The expectation is to reflect the global statistics of data noise. Indicates input data The mean prediction value output from heteroscedastic regression modeling represents the model's prediction of the input data. Expected output, Indicates input data The variance prediction values output by heteroscedastic regression modeling are used to quantify the random uncertainty of the data itself. Indicates a Gaussian distribution. This represents the observed data status of a group of sensor devices in the production environment;
[0171] For input data The cognitive uncertainty obtained by Monte Carlo sampling is used to reflect the impact of model parameters on the prediction results. Indicates the number of Monte Carlo samplings, used to assess the uncertainty of model parameters. Indicates the first The random weight matrix of the subsample. Indicates input data and random weights The model output below, This represents the expected value of the output, which is the mean of the prediction after correction for cognitive uncertainty.
[0172] In step S308 of this embodiment, the output decision analysis data is as follows:
[0173]
[0174] In the formula, Represents the characteristic function, Indicates total uncertainty. This represents the decision threshold for adaptive adjustment during dynamic confidence level adjustment. Indicates input The initial prediction confidence output, This represents a redundancy check function.
[0175] In this embodiment of the invention, to ensure the robustness of the method, a training loss function for the TD-MIN inspection model is constructed based on the above process. for:
[0176]
[0177] wherein, denotes the task loss, denotes the denoising objective function of the noise prediction network in the diffusion model, denotes the total uncertainty, denotes the expectation function, denotes the trade-off coefficient of and respectively.
[0178] In step S4 of the embodiment of the present application, according to the decision analysis data, different anomalies, alarms and faults of the working face are divided into three priorities and correspond to three different decision scenarios including:
[0179]
[0180] wherein, denotes the set of anomaly, alarm and fault events in the production environment,
[0181] Among them, the anomaly of the safety aspect belongs to the highest priority, and the decision strategy adopted is emergency shutdown, the abnormality of the device affects the coal mining condition and belongs to the medium priority, and the decision strategy adopted is to run at a reduced speed, the sensor is abnormal but does not affect the coal mining, and belongs to the lowest priority, and the decision strategy adopted is to mark for inspection; therefore, the priority of
[0182] Based on the above determined priority, the decision strategy of the decision scenario corresponding to is:
[0183]
[0184] According to the anomaly, alarm and fault information, the corresponding priority and the corresponding decision strategy obtained by the above process, the target terminal pushing rule corresponding to the decision scenario is constructed as:
[0185]
[0186] wherein, denotes the centralized control terminal in the production environment, denotes the dispatching terminal outside the production environment, denotes the mobile end APP; for example, when the production environment is the coal mine underground, denotes the centralized control terminal of various production devices underground, denotes the dispatching terminal for regulating and controlling the production devices on the surface.
[0187] In the embodiment, according to three priorities of different abnormalities, alarms and faults, the dynamic updating equation of the decision strategy is:
[0188]
[0189] Satisfies:
[0190] In the formula, represents the state of the system at time, represents that the real-time push delay approaches zero, represents all state sets of the equipment in the production environment, represents the parameter set of the production environment, represents hierarchical logic; wherein, , represents the state of the equipment at time ; , represents the parameter of the working face at time.
[0191] Through the above complete process, corresponding inspection decisions can be given according to different equipment abnormalities, alarms and faults, and the autonomous inspection of the production environment is completed.
[0192] It should be noted that the above autonomous inspection method provided by the present application can be applied to different production scenes, such as coal mine underground, factory assembly line and chemical industry park scenes; specifically, when applied to coal mine underground, the following advantages are obtained:
[0193] (1) Improve the scene perception ability in complex environment: the space of fully mechanized coal mining face and roadway in coal mine underground is narrow and closed, the equipment is dense and dynamically advances with the mining progress, and the traditional inspection is difficult to fully capture the environment and equipment state, while the method can accurately map the dispersed data such as gas concentration, equipment vibration and temperature to the underground physical space by building a three-dimensional virtual scene and associating multi-modal sensor time series data. For example, the correlation between the temperature abnormal point of a certain coal mining machine and the surrounding gas concentration can be intuitively presented in the virtual scene, helping the inspection personnel or system to quickly locate the risk source and solving the “data island” and scene perception ambiguity problem caused by the closed space in the underground.
[0194] (2) Enhance the fault tolerance ability of high interference data: there are strong interference factors such as dust, electromagnetic interference and mechanical vibration in the coal mine, and the sensor data is easy to be missing, distorted or abnormal. The method of the application builds a TD-MIN inspection big model, and the first and second model fault tolerance mechanism modules in the model can effectively compensate and feature extraction for these abnormal data through the progressive noise injection and denoising strategy of the diffusion model. For example, when the gas sensor data of a certain section is distorted due to interference, the model can repair it based on the historical time sequence rule and the surrounding associated sensor data, avoid misjudgment caused by single data abnormality, and ensure the accuracy of data analysis.
[0195] (3) Improve the reliability of decision-making in dynamic scenarios: the working face in the coal mine is constantly advancing with mining, and the equipment working condition and environmental parameters are in dynamic change, so the data distribution is easy to drift. The hierarchical uncertainty quantification network in the decoding stage can realize historical rollback and redundancy check of feature extraction by statistically quantifying data uncertainty and combining with a dynamic confidence adjustment mechanism. For example, when the operating parameters of the coal mining machine fluctuate slightly, the model will evaluate the uncertainty degree of the fluctuation, and if it is in a low confidence interval, it will combine historical data for secondary verification to determine whether to trigger an alarm, effectively reducing the misdecision caused by data drift, and making the decision more suitable for the actual dynamic scene in the coal mine.
[0196] (4) Realize the autonomous closed loop of the whole inspection process and improve the response efficiency: the traditional coal mine inspection relies on manual on-site troubleshooting, and the response is lagging and has safety risks. The method of the application can directly push the decision-making scene and corresponding strategies to the underground control terminal or ground dispatching center through autonomous analysis and decision-making and strategy pushing. For example, when detecting that the gas concentration in a certain area exceeds the standard, the model can quickly analyze the cause and generate strategies including "stop and evacuate" and "start the ventilation equipment", and push them to the relevant terminal in real time, realizing the whole process of data collection, analysis and decision execution, greatly shortening the risk response time and ensuring the safety of underground operation.
[0197] The principles and implementation methods of the application are described in the specific embodiments, and the above examples are only used to help understand the method and core idea of the application; at the same time, for those skilled in the art, according to the idea of the application, the specific implementation and application range will be changed, and the above description should not be understood as limiting the application.
[0198] Those skilled in the art will appreciate that the embodiments described herein are presented for purposes of illustration and that the inventive principles are not limited to these particular embodiments. Other variations and modifications can be made to the embodiments without departing from the spirit and scope of the inventive principles.
Claims
1. A large model-based production environment autonomous inspection method, characterized in that, Includes the following steps: S1. Build a three-dimensional virtual scene corresponding to the production environment, and associate the real-time collected multimodal sensor time series data in the three-dimensional virtual scene; S2. Construct and train the large model framework to obtain the TD-MIN inspection large model; The TD-MIN inspection model includes an encoding stage and a decoding stage. The encoding and decoding stages are respectively equipped with a first model fault tolerance mechanism module and a second model fault tolerance mechanism module. The decoding stage also includes a hierarchical uncertainty quantification network. The first and second model fault tolerance mechanism modules, by introducing a progressive noise injection and denoising strategy based on a diffusion model, achieve feature extraction of the input data and compensate for abnormal data. The hierarchical uncertainty quantification network, by statistically analyzing the uncertainty of the input data and combining it with a dynamic confidence adjustment mechanism, achieves historical backtracking and redundancy verification of the input data feature extraction. S3. Utilize the TD-MIN inspection model to autonomously analyze and make decisions on the time-series data of multimodal sensors in the corresponding 3D virtual scene of the production environment, and output decision analysis data; S4. Based on the decision analysis data, determine the decision scenarios and their corresponding decision strategies, and push them to the corresponding target terminals to achieve autonomous inspection of the production environment; In the aforementioned TD-MIN inspection model: The first model fault tolerance mechanism module and the second model fault tolerance mechanism module have the same structure, both including an input embedding layer. The output of the input embedding layer is connected to several parallel linear layers. The output of each linear layer is connected to a fault-tolerant coding layer and an attention mechanism layer in sequence. The output of each attention mechanism layer is connected to the input of a fault-tolerant feedforward neural network. The fault-tolerant feedforward neural network outputs nonlinear transformation data. The hierarchical uncertainty quantification network includes a first UNet module, a first feature extraction layer, an uncertainty quantification layer, a skip connection layer, a second feature extraction layer, a second UNet module, and a redundancy check layer connected in sequence. The redundancy check layer outputs decision analysis data after redundancy check. In step S3, the process of autonomous analysis and decision-making using the TD-MIN large model is as follows: S301. Using sensor group data as model input data, the model input data is separately addressed by location, sensor type and point cloud data through the embedding layer to obtain coded sensor group data. The sensor group data includes time-series data of multimodal sensors and their point cloud data in a three-dimensional virtual scene; S302. In the encoding stage, after the encoded sensor group data is processed through the first fully connected layer, the dependency relationship of the encoded sensor group data within its information sequence is established through the self-attention mechanism layer, and the sequence recursion of the network structure in the encoding stage is avoided by introducing a dynamic confidence adjustment mechanism. S303. The encoded sensor group data passed through the self-attention mechanism layer is spliced together through the splicing layer, and then the gradient descent calculation is performed through the second fully connected layer and the third fully connected layer in sequence before being input into the first model fault tolerance mechanism module. S304. In the first model fault tolerance mechanism module, the input data is processed and the first diffusion model is subjected to progressive noise injection and denoising, and nonlinear changing data is output. S305. The nonlinear changing data is processed sequentially through the first residual network layer, the first standard normalization layer, the first maximum pooling layer, the fourth fully connected layer, and the first Softmax layer to obtain the output data of the encoding stage. The output data of the encoding stage includes compensation for abnormal data and feature extraction data; S306. In the decoding stage, the output data from the encoding stage is processed sequentially through multiple convolution operations and a feedforward neural network, and then input into the second model fault tolerance mechanism module for processing and progressive noise injection and denoising of the second diffusion model, outputting nonlinear changing data. S307. The nonlinear change data output by the second model fault tolerance mechanism module is processed sequentially through the second residual network layer and the second standard normalization layer, and then input into the hierarchical uncertainty quantification network. S308. In the hierarchical uncertainty quantification network, multi-granularity uncertainty modeling is performed on the input data to capture the total uncertainty of the input data at different levels, and redundancy verification is performed based on it to output decision analysis data. S309. The decision analysis data is processed sequentially through the third residual network layer, the second standard normalization layer, the second max pooling layer, the fifth fully connected layer, the second Softmax layer, and the multilayer perceptron (MLP), and then output as decision analysis data.
2. The large model-based production environment autonomous inspection method according to claim 1, wherein, In step S302, the implementation process of the dynamic confidence adjustment mechanism is as follows: S302-1. Constructing the attention weight correction formula: wherein denotes the attention mechanism output result after confidence correction, denotes the query of the sensor group data, denotes the key, denotes the value, denotes the matrix, denotes the dimension of the key vector, denotes the confidence mask vector, the superscript denotes the transpose operation, denotes the element-wise multiplication; S302-2, calculate the adaptive adjusted decision threshold ; wherein denotes the initial decision threshold, denotes the decay coefficient, denotes the time at which the total uncertainty, denotes the state output at time t. S302-3, judging whether the condition is met; If yes, trigger history rollback mechanism, use history data to correct the predicted value, calculate the final predicted value ; If not, the current prediction value is used directly ; wherein, represents the final prediction value at the time instant, represents a mixing coefficient that determines the weight of the current prediction value and the historical data, represents the length of the historical window, represents the high-confidence prediction value at the historical time instant .
3. The large model-based production environment autonomous patrol inspection method of claim 1, wherein, In the first / second model fault tolerance mechanism module, the process of progressive noise injection and denoising of the diffusion model is as follows: A1. Gaussian noise is gradually injected into the original sensor group data through a Markov chain during the diffusion process of the diffusion model. The process of injecting Gaussian noise is represented as follows: wherein represents the data of the step step data step data the conditional probability distribution of the represents the data of the generated by a Gaussian noise, generated by a Gaussian noise, generated by a Gaussian noise, represents a noise schedule parameter, represents a Gaussian distribution, represents a covariance; A2. Construct a noise prediction network based on reverse process learning to denoise the injected Gaussian noise and train it to approximate the real noise. wherein the denoising objective function of the noise prediction network is: wherein denotes the actually injected Gaussian noise, denotes the noise prediction network, denotes the original input data of the unnoisy clean data sample at time step , the original data and the expectation of the Gaussian noise , denotes the original input data of the unnoisy clean data sample; A3. Based on the denoising training process, when there is drift or missing data in the sensor group, features are recovered through a conditional inverse process. The formula for recovering features through the inverse conditional process is as follows: wherein , , , denotes the denoised data at time step , denotes the noise attenuation coefficient at time step , denotes the cumulative attenuation coefficient, denotes the noise attenuation coefficient at time step , denotes the time step index, denotes the standard Gaussian noise vector.
4. The autonomous inspection method for production environments based on a large model according to claim 1, characterized in that, In the step S308, the total uncertainty of the input data at different levels is input is; wherein, represents the expected value of the aleatoric uncertainty resulting from the heteroscedastic regression modeling of the input data represents the mean prediction value output by the heteroscedastic regression modeling of the input data represents the variance prediction value output by the heteroscedastic regression modeling of the input data represents a Gaussian distribution, represents an observation of the state of a group of sensor devices in a production environment; input data cognitive uncertainty resulting from Monte Carlo sampling evaluation, denotes the number of Monte Carlo samples, denotes the random weight matrix of the denotes the model output under input data and random weight denotes the expected value of the output. 5. The autonomous inspection method for production environments based on a large model according to claim 1, characterized in that, In step S308, the output decision analysis data is as follows: In the formula, Represents the characteristic function, Indicates total uncertainty. This represents the decision threshold for adaptive adjustment during dynamic confidence level adjustment. Indicates input The initial prediction confidence output, This represents a redundancy check function.
6. The autonomous inspection method for production environments based on a large model according to claim 1, characterized in that, The training loss function of the TD-MIN inspection model for: In the formula, Indicates mission loss. Let represent the denoising objective function of the noise prediction network in the diffusion model. Indicates total uncertainty. Represents the expectation function, They represent and The tradeoff coefficient.
7. The autonomous inspection method for production environments based on a large model according to claim 1, characterized in that, In step S4, the decision-making scenario include: In the formula, This represents a collection of anomalies, alarms, and malfunctions in the production environment. ;in, The priority decreases in that order; Decision-making strategies corresponding to the decision-making scenarios for: The target terminal push rules corresponding to the decision-making scenario for: In the formula, This refers to the centralized control terminal in the production environment. This refers to a dispatch terminal outside the production environment. Indicates mobile app; The dynamic update equation for the decision-making strategy is: satisfy: In the formula, Indicates that the system is in The state at any given moment, This indicates that the real-time push latency is close to zero. This represents the set of all device states in a production environment. A set of parameters representing the production environment. This represents an event-driven decision execution function. This indicates hierarchical logic.
Citation Information
Patent Citations
Image retrieval method based on key local information
CN116467476A
Safety decision-making method for patrol scheduling of multiple unmanned vehicles in confrontation environment
CN120355259A