AI intelligent remote control system based on automatic aggregate conveying
Through adaptive feature pyramid processing and multimodal attention mechanism, combined with dark channel defogging algorithm and improved YOLOX detection algorithm, the problems of small target missed detection and dust interference failure in traditional aggregate conveying systems in dusty environments are solved, and high-precision foreign object detection is achieved.
Patent Information
- Application Number
- CN202511102279.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-09-16
AI Technical Summary
Traditional aggregate conveying systems have problems with missed detection of small targets and failure due to dust interference in dusty environments, especially the low detection accuracy for metal foreign objects smaller than 20mm. The detection recall rate decreases when the dust concentration exceeds 50mg/m3.
An adaptive feature pyramid processing module is used to perform convolution processing on the three-level feature maps. The accuracy of small target detection is improved through dynamic weight adjustment, and the multimodal attention mechanism is combined to reduce the false alarm rate. The dark channel dehazing algorithm is used to maintain a stable recognition rate in high dust environments. It integrates a distributed image acquisition module, an adaptive fill light unit and an edge computing node, combined with a multimodal Transformer architecture and an improved YOLOX detection algorithm.
The visual detection accuracy of traditional aggregate conveying systems in dusty environments is improved, the false alarm rate is reduced, and the detection capability of small targets and the robustness of the system are enhanced.
Smart Images

Figure CN120646488A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent control of mining equipment, and in particular to an AI intelligent remote control system based on automated aggregate conveying. Background Art
[0002] In mining and aggregate conveying operations, traditional conveying systems generally use belt conveyors as core transportation equipment. The control system's object detection algorithms (such as YOLOv3 / v4) often use a fixed-scale feature fusion strategy, directly splicing feature maps at different levels through FPN (Feature Pyramid Network). This approach has two major problems when handling aggregate conveying scenarios: Missed detection of small objects: Small metal foreign objects (diameter < 20 mm) on the belt surface suffer from severe feature loss after multi-layer downsampling. Dust interference failure: Traditional feature fusion is not robust enough to high-noise scenes. When the dust concentration exceeds 50mg / m 3 When the foreign body detection recall rate drops to below 65%); Therefore, to address the above problems, an AI intelligent remote control system based on automated aggregate conveying is proposed. The three-level feature map is convolved through an adaptive feature pyramid processing module. By dynamic weight adjustment, the normalized weight is increased when detecting small targets, and the detection accuracy of small targets is improved under dust interference. The false alarm rate is reduced by adopting a multimodal attention mechanism, and the dark channel defogging algorithm maintains a stable and high recognition rate even in environments with high pollutant concentrations. This improves the visual detection accuracy of traditional aggregate conveying systems in dusty environments. Summary of the Invention
[0003] In order to overcome the problems of missed detection of small targets and failure due to dust interference caused by detection blind spots in dusty environments during daily use of traditional aggregate conveying systems.
[0004] The technical solution of the present invention is: an AI intelligent remote control system based on automated aggregate conveying, comprising: Distributed image acquisition modules, consisting of a multispectral industrial camera array, are deployed at key monitoring points on the conveyor belt and are equipped with active dust removal devices and adaptive fill light units. Adaptive feature pyramid processing module, including a dynamic weight adjustment mechanism, for nonlinear fusion of multi-scale feature maps output by the YOLOX backbone network; A multimodal Transformer architecture that integrates a visual feature encoder and a vibration / temperature sensor data decoder to enable cross-modal attention feature interaction; The control execution module includes a PID regulator, a hydraulic brake device, and an audible and visual alarm, and is used to receive detection results and trigger hierarchical control instructions; Edge computing nodes with built-in improved YOLOX detection algorithm and NVIDIA Tesla T4 GPU accelerator; Remote monitoring platform, integrated with alarm log database, used to push disposal solutions.
[0005] Preferably, when the belt starts, the speed sensor triggers the active dust removal device, compressed air is sprayed periodically, the multispectral industrial camera array synchronously collects images, the light intensity sensor adjusts the brightness of the LED array in real time, the vibration sensor group collects three-axis acceleration signals, and transmits them to the edge node after FIR filtering. The image preprocessing uses an improved dark channel defogging algorithm to defog the feature map, the adaptive feature pyramid processing module dynamically fuses multi-scale features, the 5×5 convolution layer focuses on extracting belt texture features, the Transformer encoder converts image features into 768-dimensional embedding vectors, the vibration signal is decomposed into 5 IMF components by EMD, the energy entropy is calculated, the cross-modal attention layer establishes feature associations, the temperature data participates in the calculation of the modal alignment matrix M, the decision engine comprehensively evaluates the state by visual confidence, vibration energy entropy, and temperature residual, triggers secondary braking when a metal foreign body D ≥ threshold is detected, the hydraulic system increases the pressure, the deviation control uses a fuzzy PID algorithm to calculate the correction amount, the knowledge base calls the case library to match the fault characteristics, and pushes the disposal plan.
[0006] Preferably, the adaptive feature pyramid processing module includes: Three-level parallel convolutional layer, including 3×3, 5×5 and 7×5 convolution kernels, is used to process feature maps of different resolutions respectively; Dynamic weight adjustment unit, used to calculate the feature weights of each scale of the feature map , the calculation formula is: ; in, / For the / Deep feature vector of level feature map; is the Sigmoid activation function; is the exponential function operator; is the starting index of the summation operation, indicating that the feature maps of the three parallel convolution branches (3×3, 5×5, 7×5) are weighted normalized and the three-level weighted feature maps {F1, F2, F3} are output; The feature fusion channel receives the three-level weighted feature maps {F1, F2, F3} from the dynamic weight adjustment unit and uses the atrous spatial pyramid pooling (ASPP) structure to perform cross-scale feature aggregation on the three-level weighted feature maps.
[0007] Preferably, the multimodal Transformer architecture includes: The visual encoder layer processes the image feature map using a sliding window attention mechanism; The sensor decoder layer is used to generate a time series embedding vector after performing wavelet packet decomposition on the vibration signal; The cross-modal attention module is used to Implement feature interaction, where The visual feature query vector is generated by linear projection from the visual encoder output and is used to encode the key area features of the image; is the sensor feature key vector, which is output by the sensor decoder through linear projection and is used to characterize the vibration / temperature time series characteristics; is the sensor eigenvalue vector, and Homologous, after independent linear transformation; is the modal alignment matrix, which is used to correct the spatiotemporal misalignment characteristics; is the dimension scaling factor used to stabilize gradient propagation; is the matrix transpose operator; Represents the normalization function.
[0008] Preferably, the improved YOLOX detection algorithm is specifically: The decoupled detection head introduces a channel attention mechanism, and the classification branch adopts the Focal Loss loss function; The regression branch adopts the improved CIoU Loss: ,in, Represents the intersection-over-union ratio of the predicted box and the true box; Represents the Euclidean distance between the center point of the predicted box and the true box; Represents the predicted box parameters, i.e. center coordinates + width and height; Indicates the real annotation parameters; Indicates the minimum diagonal length of the bounding box; represents a learnable parameter; Represents the aspect ratio consistency measure, which is used to constrain the shape similarity between the predicted box and the real box; A dynamic sample allocation strategy is adopted in the training phase, and the positive sample threshold decreases linearly from 0.5 to 0.3 with the training rounds.
[0009] Preferably, the distributed image acquisition module includes: Active dust removal device, specifically a pulsed airflow dust removal device, linked to the belt speed sensor, with a trigger period of T = 0.2v + 1.5 (v is the belt speed in m / s); The adaptive fill light unit is an adaptive LED fill light array that adjusts the light intensity according to the ambient illumination. Automatically adjust brightness , specifically: .
[0010] Preferably, the linkage logic of the control execution module is specifically as follows: When the foreign object size D is detected to be greater than or equal to the threshold value D_th, the secondary alarm is triggered and the hydraulic brake is activated; When the belt tear detection confidence level C ≥ 0.8, the emergency stop command is executed and the positioning marking device is started; When the deviation amount δ and speed v satisfy δ>0.1v, the PID correction controller is activated.
[0011] Preferably, the remote monitoring platform comprises: The knowledge base is used to store contingency plans and maintenance process videos for typical failure cases.
[0012] Preferably, the workflow of the edge computing node includes: S801: In the preprocessing stage, an improved dark channel prior dehazing algorithm is used: ; in, represents the image after defogging, represents the original fog image, represents the transmittance, represents the global atmospheric light, is the pixel coordinate; S802: In the feature extraction stage, depthwise separable convolution is used instead of standard convolution. S803: When applying non-maximum suppression (NMS) in the post-processing stage, set the dynamic threshold θ=0.6-0.1×(target feature map size / 100).
[0013] Preferably, the system further comprises: The vibration signal analysis module performs EMD decomposition on the acceleration sensor data and extracts the energy characteristics of the first five-order IMF components; The temperature anomaly detection unit uses an LSTM network to establish a temperature time series prediction model, and triggers an early warning when the residual exceeds 3σ.
[0014] Preferably, the deployment configuration of the system satisfies: The image acquisition frame rate f and the belt speed v satisfy f≥2v (v unit: m / s); Edge node computing delay ≤ 200ms; The network transmission adopts TSN protocol to ensure the control instruction delay is less than 50ms; The overall MTBF of the system is ≥ 10,000 hours.
[0015] Beneficial effects of the present invention: The present invention performs convolution processing on the three-level feature map through an adaptive feature pyramid processing module. Through dynamic weight adjustment, the normalized weight is increased when detecting small targets, thereby improving the detection accuracy of small targets under dust interference. The false alarm rate is reduced by adopting a multimodal attention mechanism. The dark channel defogging algorithm still maintains a stable and high recognition rate in environments with high pollutant concentrations, thereby improving the visual detection accuracy of traditional aggregate conveying systems in dusty environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 Shown is a schematic diagram of the workflow of the AI intelligent remote control system based on automated aggregate conveying of the present invention; Figure 2 What is shown is a schematic diagram of the decision tree matching process in the step of pushing the disposal solution of the AI intelligent remote control system based on automated aggregate conveying of the present invention. DETAILED DESCRIPTION
[0017] The present invention will be further described below with reference to the accompanying drawings and examples.
[0018] See also Figure 1 and Figure 2 The present invention provides an embodiment: an AI intelligent remote control system based on automated aggregate conveying, comprising: Distributed image acquisition modules, consisting of a multispectral industrial camera array, are deployed at key monitoring points on the conveyor belt and are equipped with active dust removal devices and adaptive fill light units. Adaptive feature pyramid processing module, including a dynamic weight adjustment mechanism, for nonlinear fusion of multi-scale feature maps output by the YOLOX backbone network; A multimodal Transformer architecture that integrates a visual feature encoder and a vibration / temperature sensor data decoder to enable cross-modal attention feature interaction; The control execution module includes a PID regulator, a hydraulic brake device, and an audible and visual alarm, and is used to receive detection results and trigger hierarchical control instructions; Edge computing nodes with built-in improved YOLOX detection algorithm and NVIDIA Tesla T4 GPU accelerator; Remote monitoring platform, integrated with alarm log database, used to push disposal solutions.
[0019] The adaptive feature pyramid processing module includes: Three-level parallel convolutional layer, including 3×3, 5×5 and 7×5 convolution kernels, is used to process feature maps of different resolutions respectively; Dynamic weight adjustment unit, used to calculate the feature weights of each scale of the feature map , the calculation formula is: ; in, / For the / Deep feature vector of level feature map; is the Sigmoid activation function; is the exponential function operator; is the starting index of the summation operation, indicating that the feature maps of the three parallel convolution branches (3×3, 5×5, 7×5) are weighted normalized and the three-level weighted feature maps {F1, F2, F3} are output; The feature fusion channel receives the three-level weighted feature maps {F1, F2, F3} from the dynamic weight adjustment unit and uses the atrous spatial pyramid pooling (ASPP) structure to perform cross-scale feature aggregation on the three-level weighted feature maps.
[0020] In one embodiment of the present disclosure, the adaptive feature pyramid processing module, when operating, includes the following steps: Q101: Input the three-level feature map {F1, F2, F3} output by the YOLOX backbone network: F1: high-resolution feature map (1 / 8 original image size, 512×512×256); F2: medium-resolution feature map (1 / 16 original image size, 256×256×512); F3: low-resolution feature map (1 / 32 original image size, 128×128×1024); And perform convolution processing, as shown in the following table:
[0021] Q102: Feature compression: Perform global average pooling on each branch output: ; in, For the The channel feature vector of the level feature map (C=256); H and W are the height and width of the feature map; Q103: Generating dynamic weights through nonlinear mapping: ; in, is the Sigmoid activation function, ; Indicates the Normalized weights of level features ( ); When detecting small targets, B1 Value increases → Increase; under dust interference, B3 Automatic enhancement, using a large receptive field to resist noise; Q104: Cross-scale feature fusion, including: Bilinear upsampling: B2 upsamples by 2 times to 512×512; B3 upsamples by 4 times to 512×512; Channel alignment: 1×1 convolution uniformly sets the number of channels to 256; Weighted fusion: , the dimensions are 512×512×256; Q105: Multi-branch convolution, as shown in the following table:
[0022] Dilated convolution calculation: ; in, represents the expansion rate, Represents the convolution kernel weight; Perform feature splicing and compression, where splicing is done along the channel: →1280 channels; 1×1 convolution for dimensionality reduction: →256 channels; Q106: Channel attention enhancement: ; in, : channel weight vector; weighted output: ; Q107: Spatial attention enhancement: ; in, : spatial weight matrix; final output .
[0023] The multimodal Transformer architecture includes: The visual encoder layer processes the image feature map using a sliding window attention mechanism; The sensor decoder layer is used to generate a time series embedding vector after performing wavelet packet decomposition on the vibration signal; The cross-modal attention module is used to Implement feature interaction, where The visual feature query vector is generated by linear projection from the visual encoder output and is used to encode the key area features of the image; is the sensor feature key vector, which is output by the sensor decoder through linear projection and is used to characterize the vibration / temperature time series characteristics; is the sensor eigenvalue vector, and Homologous, after independent linear transformation; is the modal alignment matrix, which is used to correct the spatiotemporal misalignment characteristics; is the dimension scaling factor used to stabilize gradient propagation; is the matrix transpose operator; Represents the normalization function.
[0024] In one embodiment of the present disclosure, the multimodal Transformer architecture, when in operation, includes the following steps: Q201: Visual feature encoding: Input: Multi-scale feature map output by the backbone network ; Processing steps: Flattening the space: , where N = H × W; Positional encoding: ; Linear projection produces query vector : ; Q202: Vibration signal processing: Input: original signal from triaxial accelerometer (T=1000, 1 second data); Processing steps: Wavelet packet decomposition: 5-layer decomposition yields 16 sub-bands, and the energy characteristics of each band are extracted as follows: (i=1,...,16); Timing Embedding: (r=10 time windows); Linear projection generates keys K and values V: ( ); Q203: Modal alignment matrix generation: ; Input: Visual Position Query , sensor time key ; Fully connected network: ;parameter: , ; Output: Alignment weights , compensate for the time and space dislocation; Q204: Attention score calculation: ; in, It represents a measure of the strength of the association between the visual space position and the sensor time window; To compensate for the temporal and spatial deviations caused by belt movement; Q205: Feature aggregation: Weighted sum: ( output weights), output dimensions ; Q206: Residual connection and normalization processing: ; Q207: Visual feature reconstruction: Space restoration: ; Convolutional refinement: , the number of output channels is restored to 256; Q208: Anomaly detection decision: Confidence calculation: ;in ; Type classification: ; K is the fault type, including tearing, foreign matter, deviation, etc.
[0025] The improved YOLOX detection algorithm is specifically as follows: The decoupled detection head introduces a channel attention mechanism, and the classification branch adopts the Focal Loss loss function; The regression branch adopts the improved CIoU Loss: ,in, Represents the intersection-over-union ratio of the predicted box and the true box; Represents the Euclidean distance between the center point of the predicted box and the true box; Represents the predicted box parameters, i.e. center coordinates + width and height; Indicates the real annotation parameters; Indicates the minimum diagonal length of the bounding box; represents a learnable parameter; Represents the aspect ratio consistency measure, which is used to constrain the shape similarity between the predicted box and the real box; A dynamic sample allocation strategy is adopted in the training phase, and the positive sample threshold decreases linearly from 0.5 to 0.3 with the training rounds.
[0026] The distributed image acquisition module includes: Active dust removal device, specifically a pulsed airflow dust removal device, linked to the belt speed sensor, with a trigger period of T = 0.2v + 1.5 (v is the belt speed in m / s); The adaptive fill light unit is an adaptive LED fill light array that adjusts the light intensity according to the ambient illumination. Automatically adjust brightness , specifically: .
[0027] The linkage logic of the control execution module is specifically as follows: When the foreign object size D is detected to be greater than or equal to the threshold value D_th, the secondary alarm is triggered and the hydraulic brake is activated; When the belt tear detection confidence level C ≥ 0.8, the emergency stop command is executed and the positioning marking device is started; When the deviation amount δ and speed v satisfy δ>0.1v, the PID correction controller is activated.
[0028] In one embodiment of the present disclosure, when the control execution module is working, the specific steps are as follows: Q301: Input: Abnormality type (tear, foreign matter, deviation, etc.); Confidence ; Geometric parameters (foreign body size D, deviation δ); Position coordinates (x, y); Among them, the classification logic is: When the foreign object size D is detected to be greater than or equal to the threshold value D_th, the secondary alarm is triggered and the hydraulic brake is activated; When the belt tear detection confidence level C ≥ 0.8, the emergency stop command is executed and the positioning marking device is started; When the deviation δ and speed v satisfy δ>0.1v, the PID correction controller is activated; Q302: The emergency stop control (tear / major foreign body) execution action is: Full pressure braking of the hydraulic system; The fluorescent marker sprays marks on the defective location; The sound and light alarm activates a high-frequency alarm; The delay requirement is from signal reception to pressure peak ≤ 150ms; Q303: Secondary braking (medium-sized foreign objects): Pressure curve: (0≤t≤200ms); in, , ; Slow down the belt to 30% of the rated speed within 2 seconds; Q304: PID deviation correction (deviation): Error calculation: ; Control quantity output: ; in, (Proportional gain adapts to speed); (dynamic attenuation of integral gain); (Differential gain enhances high-speed stability); Q305: Hydraulic system control: Pressure-displacement conversion: (A is the piston area, k is the stiffness coefficient); Response from the correction agency: The hydraulic cylinder pushes the correction roller, and the correction amount Δx and the deviation amount δ are controlled in a closed loop; Position feedback sensor accuracy: ±0.5mm; Q306: Real-time monitoring of pressure sensor data and the deviation amount after correction ; The remote monitoring platform includes: The knowledge base is used to store contingency plans and maintenance process videos for typical failure cases.
[0029] In one embodiment of the present disclosure, when the remote monitoring platform is in operation, the specific steps are as follows: Q401: Feature similarity calculation: ; Where Q: current fault feature vector (128 dimensions); D: knowledge base case feature vector; Q402: Disposal plan push, decision tree matching specific steps are as follows: A1: Determine whether the vibration energy is greater than 1.2; if yes, proceed to A2; if no, proceed to A3; A2: Determine whether the temperature gradient is greater than 55°C / m; if yes, proceed to A4; if no, proceed to A5; A3: Routine checkup; A4: Emergency shutdown; A5: Partial maintenance.
[0030] The workflow of the edge computing node includes: S801: In the preprocessing stage, an improved dark channel prior dehazing algorithm is used: ; in, represents the image after defogging, represents the original fog image, represents the transmittance, represents the global atmospheric light, is the pixel coordinate; S802: In the feature extraction stage, depthwise separable convolution is used instead of standard convolution. S803: When applying non-maximum suppression (NMS) in the post-processing stage, set the dynamic threshold θ=0.6-0.1×(target feature map size / 100).
[0031] The system also includes: The vibration signal analysis module performs EMD decomposition on the acceleration sensor data and extracts the energy characteristics of the first five-order IMF components; The temperature anomaly detection unit uses an LSTM network to establish a temperature time series prediction model, and triggers an early warning when the residual exceeds 3σ.
[0032] The deployment configuration of the system meets the following requirements: The image acquisition frame rate f and the belt speed v satisfy f≥2v (v unit: m / s); Edge node computing delay ≤ 200ms; The network transmission adopts TSN protocol to ensure the control instruction delay is less than 50ms; The overall MTBF of the system is ≥ 10,000 hours.
[0033] During operation, when the belt starts, the speed sensor triggers the active dust removal device, periodically injects compressed air, the multi-spectral industrial camera array synchronously collects images, the light intensity sensor adjusts the brightness of the LED array in real time, and the vibration sensor group collects three-axis acceleration signals, which are then transmitted to the edge node after FIR filtering; Image preprocessing uses an improved dark channel dehazing algorithm to dehaze feature maps. The adaptive feature pyramid processing module dynamically fuses multi-scale features. The 5×5 convolutional layer focuses on extracting belt texture features. The Transformer encoder converts image features into a 768-dimensional embedding vector. The vibration signal is decomposed into five IMF components using EMD, and energy entropy is calculated. A cross-modal attention layer establishes feature associations. Temperature data is used to calculate the modal alignment matrix M. The decision engine integrates visual confidence, vibration energy entropy, and temperature residuals for state assessment. When a metal foreign object with a value of D ≥ the threshold is detected, secondary braking is triggered, the hydraulic system increases the pressure, and the deviation control uses a fuzzy PID algorithm to calculate the correction value. The knowledge base calls the case library to match the fault characteristics and pushes the disposal plan.
[0034] The embodiments of the present invention are described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Various changes can be made within the scope of knowledge of those skilled in the art without departing from the spirit of the present invention.
Claims
1. An AI intelligent remote control system based on automated aggregate conveying, characterized by: Includes: Distributed image acquisition modules, consisting of a multispectral industrial camera array, are deployed at key monitoring points on the conveyor belt and are equipped with active dust removal devices and adaptive fill light units. Adaptive feature pyramid processing module, including a dynamic weight adjustment mechanism, for nonlinear fusion of multi-scale feature maps output by the YOLOX backbone network; A multimodal Transformer architecture that integrates a visual feature encoder and a vibration / temperature sensor data decoder to enable cross-modal attention feature interaction; The control execution module includes a PID regulator, a hydraulic brake device, and an audible and visual alarm, and is used to receive detection results and trigger hierarchical control instructions; Edge computing nodes with built-in improved YOLOX detection algorithm and NVIDIA Tesla T4 GPU accelerator; Remote monitoring platform, integrated with alarm log database, used to push disposal solutions.
2. The AI intelligent remote control system based on automated aggregate conveying according to claim 1 is characterized in that: The adaptive feature pyramid processing module includes: Three-level parallel convolutional layer, including 3×3, 5×5 and 7×5 convolution kernels, is used to process feature maps of different resolutions respectively; Dynamic weight adjustment unit, used to calculate the feature weights of each scale of the feature map , the calculation formula is: ; in, / For the / Deep feature vector of level feature map; is the Sigmoid activation function; is the exponential function operator; is the starting index of the summation operation, indicating that the feature maps of the three parallel convolution branches (3×3, 5×5, 7×5) are weighted normalized and the three-level weighted feature maps {F1, F2, F3} are output; The feature fusion channel receives the three-level weighted feature maps {F1, F2, F3} from the dynamic weight adjustment unit and uses the atrous spatial pyramid pooling (ASPP) structure to perform cross-scale feature aggregation on the three-level weighted feature maps.
3. The AI intelligent remote control system based on automated aggregate conveying according to claim 1 is characterized in that: The multimodal Transformer architecture includes: The visual encoder layer processes the image feature map using a sliding window attention mechanism; The sensor decoder layer is used to generate a time series embedding vector after performing wavelet packet decomposition on the vibration signal; The cross-modal attention module is used to Implement feature interaction, where The visual feature query vector is generated by linear projection from the visual encoder output and is used to encode the key area features of the image; is the sensor feature key vector, which is output by the sensor decoder through linear projection and is used to characterize the vibration / temperature time series characteristics; is the sensor eigenvalue vector, and Homologous, after independent linear transformation; is the modal alignment matrix, which is used to correct the spatiotemporal misalignment characteristics; is the dimension scaling factor used to stabilize gradient propagation; is the matrix transpose operator; Represents the normalization function.
4. The AI intelligent remote control system based on automated aggregate conveying according to claim 1 is characterized in that: The improved YOLOX detection algorithm is specifically as follows: The decoupled detection head introduces a channel attention mechanism, and the classification branch adopts the Focal Loss loss function; The regression branch adopts the improved CIoU Loss: ,in, Represents the intersection-over-union ratio of the predicted box and the true box; Represents the Euclidean distance between the center point of the predicted box and the true box; Represents the predicted box parameters, i.e. center coordinates + width and height; Indicates the real annotation parameters; Indicates the minimum diagonal length of the bounding box; represents a learnable parameter; Represents the aspect ratio consistency measure, which is used to constrain the shape similarity between the predicted box and the real box; A dynamic sample allocation strategy is adopted in the training phase, and the positive sample threshold decreases linearly from 0.5 to 0.3 with the training rounds.
5. The AI intelligent remote control system based on automated aggregate conveying according to claim 1 is characterized in that: The distributed image acquisition module includes: Active dust removal device, specifically a pulsed airflow dust removal device, linked to the belt speed sensor, with a trigger period of T = 0.2v + 1.5 (v is the belt speed in m / s); The adaptive fill light unit is an adaptive LED fill light array that adjusts the light intensity according to the ambient illumination. Automatically adjust brightness , specifically: 。 6. The AI intelligent remote control system based on automated aggregate conveying according to claim 1 is characterized in that: The linkage logic of the control execution module is specifically as follows: When the foreign object size D is detected to be greater than or equal to the threshold value D_th, the secondary alarm is triggered and the hydraulic brake is activated; When the belt tear detection confidence level C ≥ 0.8, the emergency stop command is executed and the positioning marking device is started; When the deviation amount δ and speed v satisfy δ>0.1v, the PID correction controller is activated.
7. The AI intelligent remote control system based on automated aggregate conveying according to claim 1 is characterized in that: The remote monitoring platform includes: The knowledge base is used to store contingency plans and maintenance process videos for typical failure cases.
8. The AI intelligent remote control system based on automated aggregate conveying according to claim 1 is characterized in that: The workflow of the edge computing node includes: S801: In the preprocessing stage, an improved dark channel prior dehazing algorithm is used: ; in, represents the image after defogging, represents the original fog image, represents the transmittance, represents the global atmospheric light, is the pixel coordinate; S802: In the feature extraction stage, depthwise separable convolution is used instead of standard convolution. S803: When applying non-maximum suppression (NMS) in the post-processing stage, set the dynamic threshold θ=0.6-0.1×(target feature map size / 100).
9. The AI intelligent remote control system based on automated aggregate conveying according to claim 1, characterized in that: Also included are: The vibration signal analysis module performs EMD decomposition on the acceleration sensor data and extracts the energy characteristics of the first five-order IMF components; The temperature anomaly detection unit uses an LSTM network to establish a temperature time series prediction model, and triggers an early warning when the residual exceeds 3σ.
10. The AI intelligent remote control system based on automated aggregate conveying according to claim 1, characterized in that: The deployment configuration of the system meets the following requirements: The image acquisition frame rate f and the belt speed v satisfy f≥2v (v unit: m / s); Edge node computing delay ≤ 200ms; The network transmission adopts TSN protocol to ensure the control instruction delay is less than 50ms; The overall MTBF of the system is ≥ 10,000 hours.