Vision-based belt deviation intelligent detection method and system
By combining heterogeneous vision sensors and digital state simulators, the problem of early warning and in-depth diagnosis of belt misalignment detection in complex environments is solved, realizing early warning and root cause diagnosis of belt misalignment, and improving the robustness and maintenance efficiency of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CCCC MECHANICAL & ELECTRICAL ENG
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-28
AI Technical Summary
Existing belt misalignment detection technologies have weak anti-interference capabilities in complex industrial environments, cannot provide early warning of minor faults, lack in-depth diagnostic capabilities, and are complex to deploy and debug.
The system employs heterogeneous vision sensors to synchronously acquire image frame sequences and event streams, and combines them with a digital state simulator based on multibody dynamics and tribology models. By fusing multi-source data, it performs high-fidelity deduction and root cause diagnosis, generating an intelligent detection report.
It enables early warning of belt misalignment, provides in-depth diagnostic information, improves the system's robustness and maintenance efficiency under complex working conditions, and ensures the safe and stable operation of the production system.
Smart Images

Figure CN121929495A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of belt misalignment detection technology, and in particular to a vision-based intelligent method and system for detecting belt misalignment. Background Technology
[0002] In industries such as coal mining, mining, ports, and logistics, belt conveyors are core equipment for continuous material transport. During long-term, high-load operation, belts often experience lateral misalignment due to installation errors, uneven load distribution, idler roller malfunctions, or belt wear. Slight misalignment can exacerbate belt edge wear and reduce equipment lifespan; severe misalignment can lead to material spillage, belt tearing, or even equipment shutdown, causing production interruptions and safety accidents. Therefore, real-time and accurate monitoring and early warning of belt misalignment are crucial for ensuring production safety and efficiency.
[0003] Existing belt misalignment detection technologies can be mainly classified into the following categories:
[0004] 1. Detection methods based on contact sensors: This type of method involves installing mechanical limit switches, photoelectric sensors, or belt misalignment switches on both sides of the belt conveyor. When the belt deviates to the point of touching the switch, an alarm is triggered. Although this method is simple and direct, it is a reactive "after-the-fact" response, unable to provide early warning of belt misalignment trends. Furthermore, the mechanical structure is prone to wear and tear, has poor reliability, and requires significant installation and maintenance work.
[0005] 2. Image Processing Methods Based on Traditional Machine Vision: With the decrease in the cost of industrial cameras and the improvement of computing power, vision-based non-contact detection methods have become a research and application hotspot. These methods typically use fixed cameras to acquire images of belt movement, and then use image processing algorithms (such as edge detection and line fitting) to identify the belt edge position and calculate its offset. For example, Chinese invention patent application CN116573366A discloses a "Vision-Based Belt Misalignment Detection Method, System, Device, and Storage Medium." This method first acquires a template image and delineates a safe area, then performs viewpoint correction on the real-time image. Its core technology is to use the Sobel operator to perform edge detection on the image, extract multiple edge contours, and then filter the belt edge contours according to preset conditions (such as the minimum bounding rectangle size of the contour). A straight line for the belt edge is obtained through linear fitting, and finally compared with the safe area to determine whether there is misalignment. However, this type of method based on traditional image processing has the following inherent limitations:
[0006] Weak anti-interference capability: It heavily relies on clear image quality. In real industrial scenarios, factors such as changes in lighting, dust adhesion, water vapor interference, and cluttered backgrounds can significantly reduce image quality, causing edge detection algorithms to fail or generate a large amount of noise, resulting in false alarms or missed alarms.
[0007] Unable to detect early and transient anomalies: Traditional vision sensors (CMOS / CCD cameras) output complete image frames at a fixed frame rate, and their processing logic is based on "static" or "quasi-static" image analysis. For early microscopic transient events that may cause belt misalignment during belt operation (such as initial pitting vibration of idler roller bearings or momentary weak slippage of the belt), traditional cameras may fail to capture them due to motion blur or miss them due to frame rate limitations, thus failing to achieve true "early warning".
[0008] Relying on manual prior knowledge and threshold setting: This method requires manual pre-setting of "safe areas" and "preset conditions" (such as the aspect ratio of rectangles) for screening contours. These parameters need to be readjusted under different installation scenarios and belt specifications, resulting in poor generalization ability and complex deployment and debugging.
[0009] Lack of in-depth diagnostic capabilities: This method can only determine "whether the track is misaligned" and "the amount of deviation," which is a description of the phenomenon. It cannot provide root cause diagnosis for the more critical question of "why the track is misaligned" (such as which idler roller is faulty and how the tension is unbalanced), which is not conducive to preventive maintenance.
[0010] 3. Deep Learning-Based Detection Methods: In recent years, some studies have attempted to directly identify belt positions using object detection or segmentation networks. While these methods improve robustness to some extent, they require a large amount of labeled data (including images of various belt misalignment states) for model training, resulting in high data acquisition and labeling costs. Furthermore, deep learning models, being "black boxes," have difficult-to-interpret their decision-making processes, limiting their application in industrial scenarios with extremely high reliability requirements. Additionally, the computational demands of these models are typically high, posing a challenge to the real-time performance of edge devices.
[0011] In summary, existing technologies, especially the current mainstream methods based on traditional machine vision, have significant shortcomings in adaptability to complex industrial environments, early warning capabilities for minor faults, and intelligent diagnosis of the root causes of deviations. Therefore, there is an urgent need to develop a new intelligent visual inspection method and system that can adapt to complex working conditions, provide early warnings, and offer in-depth diagnostic information. Summary of the Invention
[0012] The present invention aims to address the shortcomings of the prior art by providing a vision-based intelligent detection method and system for belt misalignment.
[0013] To achieve the above objectives, the present invention adopts the following technical solution:
[0014] A vision-based intelligent method for detecting belt misalignment includes the following steps:
[0015] S1. Synchronous acquisition of heterogeneous visual data: Frame scanning vision sensors are deployed along the belt, and event cameras are deployed at key locations; the image frame sequence output by the frame scanning vision sensors and the event stream representing asynchronous changes in pixel brightness output by the event cameras are acquired synchronously.
[0016] S2. Microscopic transient visual feature extraction and early risk identification: The event stream is analyzed to identify transient visual event clusters that conform to the preset physical failure mode; when a specific event cluster is identified, an early risk event label for the target device is generated.
[0017] S3. High-fidelity simulation of belt operation status by integrating multi-source data: Based on early risk event labels, a digital status simulator corresponding to the target equipment is triggered; the digital status simulator integrates the current image frame sequence, equipment control parameters and transient visual features to perform high-frequency synchronous calculations and simulate the real-time mechanical state and spatiotemporal morphology of the belt.
[0018] S4. Simulation of Collaborative Diagnosis and Suppression Strategy for Deviation Causes Based on State Deduction: The digital state simulator, based on the deduced mechanical state, reversely solves the optimal combination of equipment parameter offsets that may lead to visually observable anomalies, as the result of deviation cause diagnosis; at the same time, it forward simulates the corrective effect of different equipment adjustment strategies on the belt state, generating a preliminary deviation suppression strategy.
[0019] S5. System-level strategy impact assessment and decision-making: The initial suppression strategy is placed in a system-level simulation model that includes related equipment and process links to conduct a simulation, assess its potential impact on the overall system operation indicators, and output globally optimized collaborative detection conclusions and suppression suggestions.
[0020] S6. Comprehensive Diagnosis and Report Generation Based on Simulation-Measurement Consistency: The belt dynamics state derived from the digital state simulator is used to generate a sequence of simulated physical quantities at the locations of non-visual physical sensors. The simulated physical quantity sequence is spatiotemporally aligned and directly compared with the corresponding non-visual physical sensor data collected during the same time period. The feature matching degree is obtained by calculating the distance or similarity between the two in a preset feature space. Based on the feature matching degree, the current confidence level of the digital state simulator is evaluated, and the credibility of the diagnostic conclusions and strategy suggestions obtained in steps S4 and S5 is dynamically calibrated accordingly. Finally, an intelligent detection report is output, which integrates visual observation, simulation deduction conclusions, and includes a comprehensive confidence level based on physical consistency.
[0021] Specifically, step S2, identifying transient visual event clusters, includes the following specific steps:
[0022] Cluster the event stream along the spatiotemporal dimension to form event clusters;
[0023] Extract the spatiotemporal features of each event cluster, including the rate of change of event density over time, the aspect ratio of the spatial distribution, and the principal axis direction;
[0024] By matching spatiotemporal features with predefined feature templates corresponding to different early failure modes, risk modes are classified and identified.
[0025] Specifically, in step S3, the digital state simulator is a physical mechanism model built based on the belt multibody dynamics and tribology model, which runs synchronously through the following hybrid architecture and method:
[0026] In terms of architecture, a hybrid architecture is adopted that couples a physical mechanism white-box model with a data-driven black-box compensation model;
[0027] During synchronization and operation: the belt edge position determined by frame scan visual data is used as a macroscopic constraint; the spatiotemporal characteristics of transient visual event clusters are used as microscopic excitations or correction inputs for specific parameters in the physical mechanism white-box model; through the state observer algorithm, the above inputs are fused to keep the internal state variables of the simulator and the actual physical state of the device with minimal deviation.
[0028] During the adaptive process, the data-driven black-box compensation model learns online based on the feature matching degree obtained in step S6, and dynamically corrects the simulation deviation of the physical mechanism white-box model under specific complex working conditions.
[0029] Specifically, in step S4, the gradient optimization algorithm is used to solve the optimal combination of equipment parameters in reverse. The equipment parameters include, but are not limited to, the equivalent friction coefficient of the idler group, the local stiffness of the belt section, and the slip rate of the drive system. The adjustment strategy of the forward simulation includes the adjustment amount of the correction roller angle, the tension force, and the drive speed.
[0030] Specifically, in step S5, the system-level simulation model reflects the topology, buffer capacity, and process logic of the material transport network; the objective function of global optimization at least considers the effectiveness of suppressing deviation, the impact on the stability of downstream equipment, and the overall energy consumption change.
[0031] Specifically, in step S6, the non-visual physical sensing data includes distributed fiber optic vibration sensing data and audio sensing data; calculating the feature matching degree between the simulation state and the physical measured data specifically includes:
[0032] Extract the spectral characteristics and simulated sound field characteristics of the simulated vibration signal output from the digital state simulator at the sensor deployment location;
[0033] The similarity between the above-mentioned simulation features and the measured spectral features and measured acoustic features collected by the corresponding sensors is calculated.
[0034] Feature matching degree is a weighted composite value of various similarity indicators, used to quantify the accuracy of simulation models in reproducing the dynamics of the physical world.
[0035] In particular, the method also includes: based on long-term statistical feature matching data, when the matching degree of a specific working condition or section is continuously lower than a preset threshold, automatically triggering an online calibration process for the corresponding physical parameters or sub-models in the digital state simulator, so as to improve its inference accuracy under that condition.
[0036] In particular, the early risk event tags generated in step S2 are not only used to trigger step S3, but are also pushed to the monitoring interface in real time for audio and visual alerts, and marked in the final intelligent detection report output in step S6 as a basis for historical tracing.
[0037] A vision-based intelligent belt misalignment detection system is used to implement a vision-based intelligent belt misalignment detection method, including:
[0038] The on-site perception layer includes:
[0039] The heterogeneous vision sensing component includes: several frame scanning vision sensors deployed on the frames on both sides of the belt conveyor to obtain panoramic images of the belt, and multiple event cameras deployed on the drive roller, the redirecting roller and the section prone to deviation.
[0040] Non-visual physical sensing components include: distributed fiber optic vibration sensors deployed along the conveyor belt frame, and an array of audio sensors deployed near key components.
[0041] The edge processing layer includes:
[0042] Edge computing gateways, deployed in the field, connect to all sensors in heterogeneous vision sensing components via industrial Ethernet or fieldbus; the edge computing gateway operates the following:
[0043] The streaming processing unit is used to receive and parse asynchronous event streams from the event camera in real time, identify transient visual event clusters that conform to preset physical failure modes, and generate early risk event tags for the target device when they are identified.
[0044] The data preprocessing and forwarding unit is used to receive and cache image frame sequences from frame-scanning vision sensors, and then package and synchronize them with early risk event tags before uploading them.
[0045] The cloud analytics layer includes:
[0046] A cloud server cluster, communicating with an industrial internet gateway and an edge computing gateway to receive visual perception data and directly communicate with non-visual physical sensing components, runs a cloud-based intelligent analysis platform; the platform includes:
[0047] The digital state simulation engine is configured to be triggered in response to early risk event tags. It is used to fuse image frame sequences, equipment control parameters and transient visual features to perform high-frequency synchronous calculations to deduce the real-time mechanical state and spatiotemporal morphology of the belt. Based on the deduced mechanical state, it reverse-solves the causes of belt deviation and forward simulation suppression strategies.
[0048] The system collaborative decision engine is configured to receive the initial suppression strategy, deduce its global impact in a system-level simulation model that includes related equipment and process links, and output optimized collaborative detection conclusions and suppression suggestions.
[0049] The model validation and report generation engine interacts with the digital state simulation engine and the system collaborative decision engine, and accesses real-time and historical data from non-visual physical sensing components. It is configured to compare the simulated physical quantity sequence derived by the digital state simulation engine with the measured data from non-visual physical sensors to calculate the feature matching degree. Based on this, it evaluates the confidence level of the deduction, the reliability of dynamic calibration diagnosis and strategy, and generates an intelligent detection report with a comprehensive confidence level.
[0050] The system interacts with data through a unified industrial data communication network, which is built on a time-series database and message middleware to support full data synchronization, storage and command transmission between the field perception layer, edge processing layer and cloud analysis layer.
[0051] The beneficial effects of this invention are:
[0052] 1. By introducing an event camera as a dynamic visual sensor, the system can capture millisecond-level microsecond transient events (such as high-frequency micro-vibrations caused by early pitting corrosion of idler roller bearings, or instantaneous local slippage of the belt) that traditional frame-scanning cameras cannot detect due to motion blur, drastic changes in lighting, or frame rate limitations. This allows the system to identify early signs of equipment deterioration before the belt exhibits visible macroscopic misalignment, achieving true predictive early warning and providing maintenance personnel with valuable intervention time to prevent the fault from escalating.
[0053] 2. A digital state simulator is employed, whose hybrid architecture integrates a white-box model of physical mechanisms based on multibody dynamics and tribology with a data-driven black-box compensation model. This not only enables high-fidelity simulation of belt operation status but also allows for reverse engineering, inferring observed visual anomalies into specific physical parameter deviations (such as abnormal friction coefficients of specific idler groups or changes in local belt stiffness). The output has been upgraded from a simple "belt misalignment alarm" to a "root cause diagnosis report," significantly improving the targeting and efficiency of maintenance work.
[0054] 3. By directly comparing simulated physical quantities (such as vibration spectra) derived from a digital state simulator with measured data collected by distributed fiber optic vibration sensors and audio sensor arrays, the system calculates the feature matching degree and can dynamically assess and calibrate the confidence level of its diagnostic conclusions. This closed-loop verification mechanism effectively overcomes the uncertainties of purely visual or purely model-based methods under complex operating conditions. Simultaneously, long-term feature matching degree data drives the model to perform online self-calibration and evolution, enabling the system to adapt to slow drifts such as equipment aging and environmental changes, maintaining high accuracy over the long term.
[0055] 4. The system adopts a layered intelligent deployment: the edge layer is responsible for processing high-throughput, low-latency event streams to achieve real-time early warning; the cloud layer aggregates multimodal data and runs complex digital twin simulations and global optimization algorithms. This architecture ensures both the real-time nature of early risk response and supports the computational needs of deep diagnostics and strategy simulation. Simultaneously, the collaborative perception and fusion of multimodal information such as vision, vibration, and acoustics significantly improves the system's overall robustness and fault tolerance in harsh industrial environments such as dust and uneven lighting.
[0056] 5. Through a system-level collaborative decision-making engine, this invention not only targets the generation suppression strategy for a single conveyor belt, but also evaluates the global impact of this strategy on upstream and downstream equipment, system throughput efficiency, and overall energy consumption within a simulation model that reflects the topology and process logic of the entire material transportation network. This makes the output strategy no longer an isolated calibration command, but a collaborative suppression scheme that has undergone global optimization and trade-offs, which helps to achieve the comprehensive goal of safe, stable, and efficient operation of the production system. Attached Figure Description
[0057] Figure 1 This is a flowchart of the method of the present invention;
[0058] Figure 2 This is a system architecture diagram of the present invention;
[0059] The following will describe in detail, with reference to the accompanying drawings, embodiments of the present invention. Detailed Implementation
[0060] The present invention will be further described below with reference to embodiments:
[0061] like Figure 1 As shown, a vision-based intelligent belt misalignment detection method includes the following steps:
[0062] S1. Heterogeneous Visual Data Synchronous Acquisition: Frame-scanning vision sensors (such as frame-scanning industrial cameras) are deployed along the conveyor belt, while event cameras are deployed at key locations (drive rollers, redirecting rollers, and sections prone to deviation). The system synchronously acquires the image frame sequences output by the frame-scanning vision sensors and the event streams representing asynchronous changes in pixel brightness output by the event cameras. Through precise hardware synchronization or software timestamp alignment technology, the image frame sequences output by the frame-scanning cameras and the asynchronous event streams output by the event cameras are strictly synchronized in time.
[0063] This step establishes a stereoscopic vision perception system that combines "macroscopic continuous observation" with "microscopic transient capture." Traditional single-frame scanning vision sensors can only provide periodic, complete scene images, which are prone to blurring or overexposure during high-speed motion or sudden changes in lighting. This step overcomes this deficiency by synchronously deploying event cameras. The event cameras output asynchronous event streams with microsecond-level latency and ultra-high dynamic range, specifically capturing instantaneous changes in pixel brightness. The combination of these two features enables the system to simultaneously possess stable monitoring capabilities of the overall belt contour, as well as sensitive capture capabilities of high-speed, subtle events such as roller vibrations and belt slippage, laying an irreplaceable data foundation for subsequent early warning.
[0064] S2. Microscopic Transient Visual Feature Extraction and Early Risk Identification: The event stream is analyzed to identify transient visual event clusters that conform to preset physical failure modes. When a specific event cluster is identified, an early risk event label is generated for the target device. Identifying transient visual event clusters includes the following specific steps:
[0065] Cluster the event stream along the spatiotemporal dimension to form event clusters;
[0066] Extract the spatiotemporal features of each event cluster, including the rate of change of event density over time (reflecting the frequency characteristics of vibration), the aspect ratio of spatial distribution, and the principal axis direction (reflecting the spatial morphology and orientation of the anomalous source).
[0067] By matching spatiotemporal features with predefined feature templates corresponding to different early failure modes, risk modes are classified and identified.
[0068] The transient visual event clusters that conform to the preset physical failure modes include at least: periodic high-frequency micro-vibration event clusters caused by early pitting corrosion of idler roller bearings; local event flow direction reversal and density drop caused by instantaneous local slippage of belts; and random spatiotemporal isolated pulse events caused by small material splashes or broken wire rope filaments.
[0069] This step enables digital feature extraction and pattern recognition of early, inherent equipment faults, significantly advancing the warning point. Traditional image edge detection-based methods can only reflect the "result" of belt position, while this step directly analyzes the "causal signals" characterizing changes in the internal state of the equipment. By performing spatiotemporal clustering on the event stream and extracting features such as density change rate and spatial distribution, abstract pixel events can be transformed into fault modes with clear physical meaning (such as "periodic high-frequency micro-vibration" corresponding to bearing pitting). This is equivalent to installing a "visual stethoscope" on the belt system, enabling the identification of potential degradation patterns and the generation of early risk event labels before a fault causes visible belt misalignment, achieving a fundamental shift from "post-event alarm" to "pre-event warning."
[0070] S3. High-fidelity simulation of belt operation status by integrating multi-source data: Based on early risk event labels, a digital status simulator corresponding to the target equipment is triggered; the digital status simulator integrates the current image frame sequence, equipment control parameters and transient visual features to perform high-frequency synchronous calculations and simulate the real-time mechanical state and spatiotemporal morphology of the belt.
[0071] The digital state simulator is a physical mechanism model built based on the belt multibody dynamics and tribology model. It runs synchronously through the following hybrid architecture and methods:
[0072] In terms of architecture, a hybrid architecture is adopted, coupling a physical mechanism white-box model and a data-driven black-box compensation model. Specifically, the physical mechanism white-box model can be derived from the model of multibody dynamics simulation software or a self-built finite-segment belt-idler discretization model. The data-driven black-box compensation model can use a fully connected neural network or a Gaussian process regression model, with the residuals output by the former (such as the difference between the predicted vibration spectrum and the actual spectrum) as the training target for online learning.
[0073] During synchronization and operation:
[0074] The belt edge position determined by frame scan visual data is used as a macroscopic constraint; specifically, the latest image acquired by the frame scan camera in step S1 is used to calculate the precise position of the two sides of the belt in real time through image processing algorithm, which serves as a strong constraint on the macroscopic spatial shape of the belt in the simulator.
[0075] The spatiotemporal characteristics of transient visual event clusters are used as micro-excitations or correction inputs for specific parameters in the physical mechanism white-box model. Specifically, the spatiotemporal characteristics (such as the frequency and intensity of micro-tremors) of the transient visual event clusters identified in step S2 are used as external micro-excitations input into the dynamic equations, or as real-time correction quantities for specific physical parameters (such as the local friction coefficient) in the white-box model.
[0076] By employing a state observer algorithm, the aforementioned inputs are fused to minimize the deviation between the simulator's internal state variables and the actual physical state of the equipment. Specifically, state observer algorithms such as Kalman filtering and particle filtering can be used to fuse the macroscopic constraints, microscopic excitations, and real-time acquired equipment control parameters (such as motor speed and tension setpoints) to drive the digital state simulator to perform high-frequency synchronous calculations. The goal is to minimize the deviation between the simulator's internal state variables (such as displacement, velocity, and stress at various points on the belt) and the actual physical state of the belt conveyor, thereby accurately reproducing the real-time mechanical state (stress and strain distribution) and spatiotemporal morphology of the entire belt length at the current moment.
[0077] During the adaptive process, the data-driven black-box compensation model learns online based on the feature matching degree obtained in step S6, and dynamically corrects the simulation deviation of the physical mechanism white-box model under specific complex working conditions.
[0078] This step creates an adaptive "digital twin" core, achieving high-precision, real-time mapping of the physical world's state to the information space. The digital state simulator in this step employs a hybrid architecture coupling a "physical mechanism white-box model" and a "data-driven black-box compensation model." Its advantages are: the physical model provides interpretability and a generalization basis, ensuring that deductions follow the laws of mechanics; the data-driven compensation model dynamically corrects the deviation between the theoretical model and the actual equipment through online learning (based on S6 feedback), solving the problem of insufficient accuracy of pure mechanism models under complex working conditions. By fusing macroscopic visual constraints and microscopic event features through a state observer algorithm, the simulator's internal state continuously maintains minimal deviation from the real physical state, thereby outputting a reliable real-time mechanical state and three-dimensional morphology of the belt, providing a unique and authentic source for in-depth diagnostics.
[0079] S4. Simulation of Collaborative Diagnosis and Suppression Strategy for Deviation Causes Based on State Deduction: The digital state simulator, based on the deduced mechanical state, reversely solves the optimal combination of equipment parameter offsets that may lead to visually observable anomalies, as the result of deviation cause diagnosis; at the same time, it forward simulates the corrective effect of different equipment adjustment strategies on the belt state, generating a preliminary deviation suppression strategy.
[0080] The optimal combination of equipment parameters is solved in reverse using a gradient optimization algorithm. The equipment parameters include, but are not limited to, the equivalent friction coefficient of the idler group, the local stiffness of the belt section, and the slip rate of the drive system. The adjustment strategies for forward simulation include the adjustment of the correction roller angle, the tension force, and the drive speed.
[0081] Specifically, the simulated belt morphology (which can be considered the "effect") is linked to the visually observed "belt misalignment" or "abnormal vibration." Using optimization algorithms such as gradient descent, the simulator internally solves for which combinations of equipment parameter offsets (e.g., a 10% decrease in the equivalent friction coefficient of idler group 2 and a 5% increase in the local stiffness of the belt in section C) can cause the simulator to output the observed abnormal state with minimal error. The optimal parameter offset combination is the diagnostic result of the root cause of the misalignment or abnormality. After identifying the "cause," the digital state simulator can simulate various possible "treatment plans." For example, it can simulate different adjustment strategies such as adjusting the straightening roller angle by +1°, increasing the tension by 5kN, or reducing the drive speed by 2%, and quickly deduce the effect of each strategy on correcting the belt misalignment. This generates one or more preliminary, targeted misalignment suppression strategies.
[0082] This step provides causal reasoning and decision support capabilities, from "perceiving the phenomenon" to "diagnosing the root cause" and "pre-planning countermeasures." Traditional methods stop at "detecting deviation." This step utilizes a high-fidelity digital state simulator for inverse solving (such as gradient optimization). Through inverse calculation, the observed abnormal state is traced back to specific equipment parameter deviations, providing a quantitative answer to "why the deviation occurred." Simultaneously, through forward simulation, the correction effects of different adjustment strategies (such as adjusting the angle of the correction roller) can be virtually tested in advance, generating preliminary suppression strategies. This transforms maintenance from "trial and error based on experience" to "precise decision-making based on simulation," greatly improving maintenance efficiency and the first-time success rate.
[0083] S5. System-level strategy impact assessment and decision-making: The initial suppression strategy is placed in a system-level simulation model that includes related equipment and process links to conduct a simulation, assess its potential impact on the overall system operation indicators, and output globally optimized collaborative detection conclusions and suppression suggestions.
[0084] The system-level simulation model reflects the topology, buffer capacity, and process logic of the material transport network; the objective function of global optimization should at least consider the effectiveness of suppressing deviation, the impact on the stability of downstream equipment, and the overall energy consumption change.
[0085] Specifically, considering that the belt conveyor is part of a complex material transport system, local adjustments may affect upstream and downstream processes. Therefore, the preliminary suppression strategy generated in step S4 is input into a higher-level system-level simulation model. This model reflects the topology of the entire material transport network, the buffer capacity of the silos, and the process interlocking logic. The model is used to evaluate whether the preliminary suppression strategy is effective not only in correcting the machine's deviation but also in its potential impact on key operating indicators such as the stability of the downstream feeder and the overall energy consumption of the system. Through a multi-objective optimization algorithm (the objective function comprehensively considers the effectiveness of the correction, downstream impacts, and energy consumption changes), the preliminary strategy is globally optimized and balanced, ultimately outputting a globally optimized collaborative detection conclusion and suppression recommendations.
[0086] This step integrates localized solutions for single-point failures into the overall production system for collaborative optimization. Suppression strategies for a single conveyor belt can have cascading effects on downstream material supply and system energy consumption. This step uses a system-level simulation model to deduce the initial strategy in a virtual environment encompassing network topology, buffer zones, and process logic. Its advantage lies in its ability to assess the comprehensive impact of the strategy from a global optimization perspective (the objective function covers deviation correction, downstream stability, and overall energy consumption), and to output collaboratively optimized decision recommendations. This ensures that corrective measures are not only effective but also aligned with the overall production system's operational goals, achieving a shift in thinking from "equipment maintenance" to "system operation and maintenance."
[0087] S6. Comprehensive Diagnosis and Report Generation Based on Simulation-Measurement Consistency: The belt dynamics state derived from the digital state simulator is used to generate a sequence of simulated physical quantities at the locations of non-visual physical sensors. The simulated physical quantity sequence is spatiotemporally aligned and directly compared with the corresponding non-visual physical sensor data collected during the same time period. The feature matching degree is obtained by calculating the distance or similarity between the two in a preset feature space. Based on the feature matching degree, the current confidence level of the digital state simulator is evaluated, and the credibility of the diagnostic conclusions and strategy suggestions obtained in steps S4 and S5 is dynamically calibrated accordingly. Finally, an intelligent detection report is output, which integrates visual observation and simulation deduction conclusions, and includes a comprehensive confidence level based on physical consistency.
[0088] Non-visual physical sensing data includes distributed fiber optic vibration sensing data and audio sensing data; calculating the feature matching degree between the simulation state and the physical measurement data specifically includes:
[0089] Extract the spectral characteristics and simulated sound field characteristics of the simulated vibration signal output from the digital state simulator at the sensor deployment location;
[0090] The similarity between the above-mentioned simulation features and the measured spectral features and measured acoustic features collected by the corresponding sensors is calculated.
[0091] Feature matching degree is a weighted composite value of various similarity indicators, used to quantify the accuracy of simulation models in reproducing the dynamics of the physical world.
[0092] The method also includes: based on long-term statistical feature matching data, when the matching degree of a specific working condition or section is continuously lower than a preset threshold, automatically triggering an online calibration process for the corresponding physical parameters or sub-models in the digital state simulator to improve its inference accuracy under that condition.
[0093] The early risk event tags generated in step S2 are used to trigger step S3, and are also pushed to the monitoring interface in real time for audio and visual alerts, and marked in the final intelligent detection report output in step S6 as a basis for historical tracing.
[0094] Specifically, to ensure the reliability of the entire analysis chain, this invention introduces a crucial "physical consistency verification" step. After deriving the belt dynamics, the digital state simulator can calculate, based on the physical model, the theoretically expected sequence of simulated physical quantities, such as simulated vibration signals and simulated sound field signals, at locations where non-visual physical sensors (e.g., distributed fiber optic vibration sensors, audio sensor arrays) are actually deployed. In the model verification and report generation engine, these simulated physical quantity sequences are rigorously spatiotemporally aligned and directly compared with measured physical sensing data collected from real sensors during the same time period. By calculating the distance or similarity between the simulated and measured signals in preset feature spaces such as spectral and acoustic characteristics, a quantified feature matching degree is obtained. Based on this matching degree, the confidence level of the digital state simulator's deduction under the current operating conditions can be objectively evaluated. A high confidence level indicates high credibility of the diagnostic conclusions and strategy recommendations in steps S4 and S5; a low confidence level will automatically reduce the weight of these conclusions and recommendations and will be clearly indicated in the report. Simultaneously, this matching degree is also used to drive the online learning of the data-driven black-box compensation model, dynamically correcting the biases of the white-box model. Finally, the system automatically generates an intelligent detection report. This report organically integrates: a summary of the original visual observations (images / events), the state and cause diagnosis results deduced from the digital state simulator, and system-level optimized suppression strategy recommendations. Most importantly, the report includes a comprehensive confidence level (e.g., "high / medium / low") based on physical consistency verification, making the reliability of the report content immediately clear to operators. Early risk event tags generated in step S2 are also marked in the report for easy historical tracing.
[0095] This step establishes a closed-loop "perception-simulation-verification" credibility assessment system, enabling the system to self-evaluate, self-calibrate, and continuously evolve. This step introduces independent multiphysics sensor data (vibration, sound) as "physical world truth," directly and quantitatively comparing it with the simulation predictions of the digital state simulator to generate feature matching degrees. Its advantages include: dynamic credibility calibration: dynamically adjusting the confidence level of diagnostic conclusions based on matching degrees, outputting a report with a comprehensive confidence level, significantly improving the credibility of the results; driving model self-evolution: long-term low matching degree statistics can automatically trigger online calibration of the digital state simulator, enabling the model to adapt to equipment aging and changing operating conditions; achieving multimodal fusion: through the alignment of the same physical quantity (such as vibration spectrum) in the simulation and measurement domains, deep fusion and mutual verification of evidence from vision, mechanistic models, and physical sensors are achieved, significantly enhancing the system's robustness.
[0096] like Figure 2 As shown, a vision-based intelligent belt misalignment detection system is used to implement a vision-based intelligent belt misalignment detection method, including:
[0097] The field perception layer is the data source of the system, and it includes:
[0098] A heterogeneous vision sensing component that integrates two vision sensors based on different principles. It includes: several frame-scanning vision sensors deployed on the frames on both sides of the belt conveyor to acquire panoramic images of the belt, and multiple event cameras deployed on the drive roller, redirecting roller, and sections prone to deviation, which are used to capture event streams characterized by asynchronous changes in pixel brightness caused by microphysical changes;
[0099] Non-visual physical sensing components provide physical field observation data independent of vision. These include: distributed fiber optic vibration sensors deployed along the conveyor frame (providing vibration signals across the entire line with precise location information), and an array of audio sensors deployed near key components (drive motors, large idler rollers, etc.) for collecting operating sounds and locating sound sources.
[0100] The edge processing layer is responsible for real-time, lightweight processing of on-site data, and includes:
[0101] Edge computing gateways, deployed in the field, connect to all sensors in heterogeneous vision sensing components via industrial Ethernet or fieldbus; the edge computing gateway operates the following:
[0102] A streaming processing unit receives and parses asynchronous event streams from event cameras in real time, identifies transient visual event clusters that conform to preset physical failure modes, and generates early risk event tags for target devices upon identification. In a preferred embodiment, this unit uses an optimized spatiotemporal clustering algorithm (such as density-based temporal surface clustering) to parse the event stream. Its workflow is as follows: first, the event stream is clustered within a set time window to form discrete event clusters; then, the spatiotemporal features of each event cluster are extracted, such as the rate of change of event density over time and spatial distribution morphology (aspect ratio, main axis direction); finally, these features are matched with a predefined feature template library corresponding to different early physical failure modes (such as bearing pitting, belt slippage). Once a specific event cluster conforming to a preset failure mode is identified, the unit immediately generates a structured early risk event tag. This tag includes at least an event ID, a high-precision timestamp, an associated device identifier, a failure mode type, a confidence score, and spatial location information.
[0103] The data preprocessing and forwarding unit is used to receive and cache image frame sequences from frame-scanning vision sensors, and then package and synchronize them with early risk event tags before uploading.
[0104] The cloud analytics layer is the intelligent hub of the system, and it includes:
[0105] A cloud server cluster, communicating with an industrial internet gateway and an edge computing gateway to receive visual perception data and directly communicate with non-visual physical sensing components, runs a cloud-based intelligent analysis platform; the platform includes:
[0106] The digital state simulation engine is configured to be triggered in response to early risk event tags. It fuses image frame sequences, equipment control parameters, and transient visual features to perform high-frequency synchronous calculations to deduce the real-time mechanical state and spatiotemporal morphology of the belt. Based on the deduced mechanical state, it inversely solves for the causes of belt misalignment and develops forward simulation suppression strategies. Specifically, the engine is triggered upon receiving early risk event tags from the edge. Internally, the engine employs a hybrid architecture, coupling a physical mechanism white-box model based on belt multibody dynamics and tribology principles (such as a belt-idler system model established using a finite segment discretization method) and a data-driven black-box compensation model based on a neural network. The engine fuses image frame sequences (used to extract the macroscopic edge positions of the belt as constraints), real-time equipment control parameters (tension, speed, etc.), and transient visual features. Through a state observer algorithm (such as a Kalman filter), it performs millisecond-level high-frequency synchronous calculations to deduce the real-time three-dimensional morphology, stress distribution, and vibration state of the belt, which is highly consistent with the physical world.
[0107] The system collaborative decision-making engine is configured to receive preliminary suppression strategies and extrapolate their global impact within a system-level simulation model that includes related equipment and process steps. It then outputs optimized collaborative detection conclusions and suppression recommendations. Specifically, the engine maintains a system-level simulation model that accurately reflects the topology of the entire material transport network, the buffer capacity of the storage bins, and the process interlocking logic. Within this model, the engine performs forward extrapolation of the preliminary strategies, evaluating their global impact through multi-objective optimization (the objective function comprehensively considers the effectiveness of suppression, the stability impact on downstream equipment, and changes in overall system energy consumption). Finally, it outputs collaboratively optimized final decision recommendations.
[0108] The model validation and report generation engine interacts with the digital state simulation engine and the system collaborative decision engine, and accesses real-time and historical data from non-visual physical sensing components. It is configured to compare the simulated physical quantity sequence derived by the digital state simulation engine with measured data from non-visual physical sensors to calculate the feature matching degree. Based on this, it evaluates the confidence level of the derivation, the reliability of dynamic calibration diagnosis and strategy, and generates an intelligent detection report with a comprehensive confidence level. Specifically, the similarity calculation can use cosine similarity, dynamic time warping (DTW) distance or spectral correlation coefficient.
[0109] The system interacts with data through a unified industrial data communication network, which is built on a time-series database and message middleware to support full data synchronization, storage and command transmission between the field perception layer, edge processing layer and cloud analysis layer.
[0110] The system has fault tolerance and adaptive mechanisms:
[0111] Network degradation mode: When the edge and cloud networks are interrupted, the edge computing gateway can operate independently, continuously perform event stream analysis and local early warning, cache data locally, and automatically synchronize after the network is restored.
[0112] Online Model Evolution: The model validation and report generation engine continuously monitors feature matching rates under various operating conditions. When the matching rate under a specific operating condition (such as extreme humidity or high load) is found to be consistently below a threshold, an online incremental learning process for the data-driven compensation model in the digital state simulation engine is automatically triggered to dynamically correct physical model deviations, enabling the system to continuously adapt to equipment aging and environmental changes.
[0113] In summary, this invention constructs a novel intelligent detection paradigm for belt misalignment by employing heterogeneous visual fusion, digital state simulator deduction, multi-physics field verification, and closed-loop decision optimization. This paradigm achieves a leap from "appearance perception" to "mechanism diagnosis" and then to "optimization decision-making," significantly improving the safety, reliability, and intelligence level of industrial conveying systems.
[0114] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," "counterclockwise," "axial," "radial," and "circumferential" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.
[0115] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection, an electrical connection, or a connection that allows communication between them; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components, unless otherwise explicitly limited. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0116] The present invention has been described above by way of example. Obviously, the specific implementation of the present invention is not limited to the above-described manner. Any improvements made by adopting the inventive concept and technical solution of the present invention, or direct application to other occasions without modification, are all within the protection scope of the present invention.
Claims
1. A vision-based intelligent method for detecting belt misalignment, characterized in that, Includes the following steps: S1. Synchronous acquisition of heterogeneous visual data: Frame scanning visual sensors are deployed along the belt, while event cameras are deployed at key locations. The image frame sequence output by the synchronous frame-scanning vision sensor is acquired, and the event stream representing the asynchronous changes in pixel brightness output by the event camera is acquired. S2. Microscopic transient visual feature extraction and early risk identification: Analyze the event stream and identify transient visual event clusters that conform to the preset physical failure mode; When a specific cluster of events is identified, an early risk event label is generated for the target device; S3. High-fidelity simulation of belt operation status by integrating multi-source data: Based on early risk event labels, trigger the digital status simulator corresponding to the target equipment; The digital state simulator integrates the current image frame sequence, equipment control parameters, and transient visual features to perform high-frequency synchronous calculations and deduce the real-time mechanical state and spatiotemporal morphology of the belt. S4. Simulation of Collaborative Diagnosis and Suppression Strategy for Deviation Causes Based on State Deduction: The digital state simulator, based on the deduced mechanical state, reversely solves the optimal combination of equipment parameter offsets that may lead to visually observable anomalies, as the result of deviation cause diagnosis; at the same time, it forward simulates the corrective effect of different equipment adjustment strategies on the belt state, generating a preliminary deviation suppression strategy. S5. System-level strategy impact assessment and decision-making: The initial suppression strategy is placed in a system-level simulation model that includes related equipment and process links to conduct a simulation, assess its potential impact on the overall system operation indicators, and output globally optimized collaborative detection conclusions and suppression suggestions. S6. Comprehensive diagnosis and report generation based on simulation-measurement consistency: The belt dynamic state derived by the digital state simulator is used to generate a simulated physical quantity sequence at the location of non-visual physical sensor deployment; the simulated physical quantity sequence is spatiotemporally aligned and directly compared with the corresponding non-visual physical sensor data collected in the same time period, and the feature matching degree is obtained by calculating the distance or similarity between the two in the preset feature space. Based on feature matching degree, the current inference confidence level of the digital state simulator is evaluated, and the credibility of the diagnostic conclusions and strategy recommendations obtained in steps S4 and S5 is dynamically calibrated accordingly. Finally, an intelligent detection report is output, which integrates visual observation, simulation inference conclusions, and includes a comprehensive confidence level based on physical consistency.
2. The vision-based intelligent belt misalignment detection method according to claim 1, characterized in that, In step S2, identifying transient visual event clusters includes the following specific steps: Cluster the event stream along the spatiotemporal dimension to form event clusters; Extract the spatiotemporal features of each event cluster, including the rate of change of event density over time, the aspect ratio of the spatial distribution, and the principal axis direction; By matching spatiotemporal features with predefined feature templates corresponding to different early failure modes, risk modes are classified and identified.
3. The vision-based intelligent belt misalignment detection method according to claim 1, characterized in that, In step S3, the digital state simulator is a physical mechanism model built based on the belt multibody dynamics and tribology model, which runs synchronously through the following hybrid architecture and method: In terms of architecture, a hybrid architecture is adopted that couples a physical mechanism white-box model with a data-driven black-box compensation model; During synchronization and operation: the belt edge position determined by frame scan visual data is used as a macroscopic constraint; the spatiotemporal characteristics of transient visual event clusters are used as microscopic excitations or correction inputs for specific parameters in the physical mechanism white-box model; through the state observer algorithm, the above inputs are fused to keep the internal state variables of the simulator and the actual physical state of the device with minimal deviation. During the adaptive process, the data-driven black-box compensation model learns online based on the feature matching degree obtained in step S6, and dynamically corrects the simulation deviation of the physical mechanism white-box model under specific complex working conditions.
4. The vision-based intelligent belt misalignment detection method according to claim 1, characterized in that, In step S4, the gradient optimization algorithm is used to solve the optimal combination of equipment parameters in reverse. The equipment parameters include, but are not limited to, the equivalent friction coefficient of the idler group, the local stiffness of the belt section, and the slip rate of the drive system. The adjustment strategy of the forward simulation includes the adjustment amount of the correction roller angle, the tension force, and the drive speed.
5. The vision-based intelligent belt misalignment detection method according to claim 1, characterized in that, In step S5, the system-level simulation model reflects the topology, buffer capacity, and process logic of the material transportation network; the objective function of global optimization at least considers the effectiveness of suppressing deviation, the impact on the stability of downstream equipment, and the overall energy consumption change.
6. The vision-based intelligent belt misalignment detection method according to claim 1, characterized in that, In step S6, the non-visual physical sensing data includes distributed fiber optic vibration sensing data and audio sensing data; calculating the feature matching degree between the simulation state and the physical measured data specifically includes: Extract the spectral characteristics and simulated sound field characteristics of the simulated vibration signal output from the digital state simulator at the sensor deployment location; The similarity between the above-mentioned simulation features and the measured spectral features and measured acoustic features collected by the corresponding sensors is calculated. Feature matching degree is a weighted composite value of various similarity indicators, used to quantify the accuracy of simulation models in reproducing the dynamics of the physical world.
7. The vision-based intelligent belt misalignment detection method according to claim 6, characterized in that, The method also includes: based on long-term statistical feature matching data, when the matching degree of a specific working condition or section is continuously lower than a preset threshold, automatically triggering an online calibration process for the corresponding physical parameters or sub-models in the digital state simulator to improve its inference accuracy under that condition.
8. The vision-based intelligent belt misalignment detection method according to claim 1, characterized in that, The early risk event tags generated in step S2 are used to trigger step S3, and are also pushed to the monitoring interface in real time for audio and visual alerts, and marked in the final intelligent detection report output in step S6 as a basis for historical tracing.
9. A vision-based intelligent belt misalignment detection system, used to implement the vision-based intelligent belt misalignment detection method according to any one of claims 1-8, characterized in that, include: The on-site perception layer includes: The heterogeneous vision sensing component includes: several frame scanning vision sensors deployed on the frames on both sides of the belt conveyor to obtain panoramic images of the belt, and multiple event cameras deployed on the drive roller, the redirecting roller and the section prone to deviation. Non-visual physical sensing components include: distributed fiber optic vibration sensors deployed along the conveyor belt frame, and an array of audio sensors deployed near key components. The edge processing layer includes: Edge computing gateways, deployed in the field, connect to all sensors in heterogeneous vision sensing components via industrial Ethernet or fieldbus; the edge computing gateway operates the following: The streaming processing unit is used to receive and parse asynchronous event streams from the event camera in real time, identify transient visual event clusters that conform to preset physical failure modes, and generate early risk event tags for the target device when they are identified. The data preprocessing and forwarding unit is used to receive and cache image frame sequences from frame-scanning vision sensors, and then package and synchronize them with early risk event tags before uploading them. The cloud analytics layer includes: A cloud server cluster, communicating with an industrial internet gateway and an edge computing gateway to receive visual perception data and directly communicate with non-visual physical sensing components, runs a cloud-based intelligent analysis platform; the platform includes: The digital state simulation engine is configured to be triggered in response to early risk event tags. It is used to fuse image frame sequences, equipment control parameters and transient visual features to perform high-frequency synchronous calculations to deduce the real-time mechanical state and spatiotemporal morphology of the belt. Based on the deduced mechanical state, it reverse-solves the causes of belt deviation and forward simulation suppression strategies. The system collaborative decision engine is configured to receive the initial suppression strategy, deduce its global impact in a system-level simulation model that includes related equipment and process links, and output optimized collaborative detection conclusions and suppression suggestions. The model validation and report generation engine interacts with the digital state simulation engine and the system collaborative decision engine, and accesses real-time and historical data from non-visual physical sensing components. It is configured to compare the simulated physical quantity sequence derived by the digital state simulation engine with the measured data from non-visual physical sensors to calculate the feature matching degree. Based on this, it evaluates the confidence level of the deduction, the reliability of dynamic calibration diagnosis and strategy, and generates an intelligent detection report with a comprehensive confidence level. The system interacts with data through a unified industrial data communication network, which is built on a time-series database and message middleware to support full data synchronization, storage and command transmission between the field perception layer, edge processing layer and cloud analysis layer.
Citation Information
Patent Citations
Vision-based belt deviation detection method, system and equipment and storage medium
CN116573366A