Winch intelligent monitoring system based on machine vision

By using dynamic geometric consistency modeling and multimodal feature fusion technology, the problem of insufficient real-time performance in the winch monitoring system was solved, enabling accurate identification and adaptive early warning of the winch's operating status, and improving the intelligence level and safety of the monitoring system.

CN121564643AInactive Publication Date: 2026-02-24合肥瑞徽人工智能研究院有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511669999.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-14
Publication Date
2026-02-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing hoist monitoring systems rely on manual inspections or single electromechanical sensors, making it difficult to identify early faults in real time. Furthermore, the lack of multimodal fusion between visual information and electromechanical signals results in a single detection dimension, insufficient real-time performance, and an inability to accurately characterize the dynamic features of the mechanical structure.

Method used

A machine vision-based intelligent monitoring system for winches is adopted, which integrates dynamic geometric consistency modeling, electromechanical signal feature analysis and multimodal feature fusion technology. Through visual acquisition, electromechanical signal acquisition, time synchronization and data preprocessing, dynamic attitude estimation, multimodal feature fusion and intelligent recognition, the system can accurately identify the operating status of the winch and provide adaptive early warning.

Benefits of technology

It improves monitoring accuracy and response speed, enhances environmental adaptability, maintains consistent identification results under complex working conditions, achieves accurate identification of winch operating status and fault early warning, and improves the intelligence level and safety protection capabilities of monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121564643A_ABST
    Figure CN121564643A_ABST
Patent Text Reader

Abstract

The invention discloses a winch intelligent monitoring system based on machine vision, and the system comprises a vision collection module which is used for obtaining the images of a winding drum and a steel wire rope; the electromechanical signal acquisition module is used for acquiring current, rotating speed, tension, vibration and attitude angle signals; the time synchronization and preprocessing module is used for aligning and normalizing the multi-source data; the dynamic attitude estimation module is used for calculating rotation and translation attitudes and extracting visual features; the adaptive feature modeling module is used for keeping geometric consistency and dynamically adjusting constraint weights; the multi-modal fusion module is used for fusing visual and electromechanical features; the intelligent identification module is used for identifying an operation state and a fault type; and the early warning control module is used for outputting graded early warning and linkage control. According to the invention, intelligent identification and dynamic early warning of the operation state of the winch are realized, and the monitoring precision, safety and reliability are obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of machine vision monitoring, electromechanical signal analysis and intelligent control technology, and in particular to an intelligent monitoring system for winches based on machine vision. Background Technology

[0002] As a critical lifting device widely used in mining, port loading and unloading, construction, and shipping, the safety and stability of winches directly affect the efficiency of engineering operations and the safety of personnel. Currently, winch monitoring mainly relies on manual inspections or data collection from single electromechanical sensors, such as parameters like speed, current, tension, vibration, and temperature. While these traditional methods can reflect some operational statuses, they suffer from limitations such as limited detection dimensions, insufficient real-time performance, and inability to accurately characterize the dynamic features of the mechanical structure. When early signs of problems arise, such as wire rope wear, drum sway, abnormal bearing vibration, or motor overload, single sensor signals often fail to identify them in a timely manner, leading to delayed fault warnings.

[0003] With the development of machine vision and intelligent sensing technologies, image recognition-based equipment monitoring is increasingly being applied in the field of lifting machinery. However, winches experience continuous rotation, vibration, and changes in ambient lighting during operation, making it difficult for traditional vision algorithms to maintain stable identification of the drum and wire rope, and feature extraction is easily affected by posture changes. Furthermore, existing research often analyzes visual information and electromechanical signals independently, lacking an effective multimodal fusion mechanism, thus failing to achieve a unified description and intelligent judgment of the equipment's operating status.

[0004] Therefore, there is an urgent need for an intelligent monitoring system for winches that can integrate machine vision and electromechanical signal features. The system should possess the ability to dynamically model the winch's rotational attitude, maintaining consistent identification results under different viewing angles, vibrations, and load variations. Simultaneously, through multi-source feature joint embedding and risk classification strategies, it should achieve accurate identification and adaptive early warning of the winch's operating status, anomaly types, and fault locations, thereby effectively improving the intelligence level and safety protection capabilities of winch monitoring. Summary of the Invention

[0005] One objective of this invention is to propose an intelligent monitoring system for winches based on machine vision. This invention integrates dynamic geometric consistency modeling, electromechanical signal feature analysis, and multimodal feature fusion technology to comprehensively describe the winch's operating status from multi-source data acquisition, attitude estimation, feature alignment to risk warning and coordinated control. The system achieves stable recognition of visual features under attitude changes through a dynamic geometric constraint mechanism and constructs a multimodal joint representation by combining electromechanical signal features, forming a four-level risk warning and control strategy. This invention has the advantages of high monitoring accuracy, fast response speed, and strong environmental adaptability, and can be used for intelligent monitoring and fault prevention of winches.

[0006] According to an embodiment of the present invention, a machine vision-based intelligent monitoring system for winches includes:

[0007] The visual acquisition module acquires continuous image sequences of the winch drum and wire rope during operation, generating visual data. The electromechanical signal acquisition module synchronously acquires raw electromechanical signals of the winch, including current, speed, tension, vibration, and attitude angles. The time synchronization and data preprocessing module performs time alignment, filtering, and normalization on the visual data and electromechanical signals, generating a multi-source data sequence. The dynamic attitude estimation module calculates the winch's rotation matrix and translation vector in real time based on changes in angular velocity, tilt angle, and visual feature points in the multi-source data sequence, representing its spatial attitude; it also extracts visual features from the visual data to provide input for subsequent modeling. Based on dynamic geometry... The adaptive feature modeling mechanism with geometric consistency constraints is used to maintain visual feature consistency under attitude change conditions. It performs spatial alignment based on rotation matrix and translation vector, and dynamically adjusts constraint weights according to the operating status. The multimodal feature fusion module fuses visual features processed by geometric consistency constraints with electromechanical signal features to generate multimodal feature vectors reflecting the operating status of the winch. The intelligent recognition module identifies the operating status, anomaly type, and fault location of the winch based on the multimodal feature vectors, forming a comprehensive recognition result set. The early warning and control module outputs risk warnings based on the comprehensive recognition result set and controls the winch to decelerate or stop in the event of a serious anomaly.

[0008] Optionally, modules can be integrated using the following methods:

[0009] S1. Acquire a continuous image sequence of the winch drum and wire rope during operation to generate visual data; S2. Synchronously acquire raw electromechanical signals of current, rotation speed, tension, vibration, and attitude angle; S3. Perform time alignment, filtering, and normalization on the visual data and raw electromechanical signals to generate a multi-source data sequence; S4. Based on the changes in angular velocity, tilt angle, and visual feature points of adjacent frames in the multi-source data sequence, calculate the rotation matrix and translation vector in real time; extract the visual features at the corresponding time from the visual data;

[0010] S5. An adaptive feature modeling mechanism based on dynamic geometric consistency constraints spatially aligns visual features according to the rotation matrix and translation vector, and dynamically adjusts the geometric constraint weights based on the electromechanical signal feature parameters of speed, vibration, current, and attitude angle provided by the electromechanical signal acquisition module and the time synchronization and data preprocessing module; S6. The visual features and electromechanical signal features are jointly embedded and fused to generate a multimodal feature vector reflecting the overall operating status of the winch; the electromechanical signal features are feature results extracted based on the original electromechanical signals, including mean, variance, frequency domain energy, or time-series rate of change; S7. Classification and regression analysis are performed based on the multimodal feature vector, and the operating status identification result, anomaly type identification result, and fault location identification result of the winch are output respectively, forming a comprehensive identification result set; S8. A risk classification early warning signal is output based on the comprehensive identification result set, and the winch is controlled to execute deceleration or shutdown commands in the event of a serious anomaly.

[0011] Optionally, S4 specifically includes:

[0012] S41. Select a data window for the current time and adjacent times from a multi-source data sequence according to a unified timestamp;

[0013] S42. Detect corner points or edge visual feature points on adjacent image frames respectively, and calculate the correspondence between adjacent frames based on similarity measurement to form a matching result set composed of visual feature points; use consistency test method to remove matching points that do not meet geometric constraints to obtain a set of usable feature point correspondences.

[0014] S43. Based on the device intrinsic parameters of the visual acquisition module, the feature points in the matching result set are normalized, and the initial relative pose values ​​between adjacent image frames are calculated based on the normalized feature points to obtain preliminary rotation parameters and translation direction parameters; S44. The angular velocity is numerically integrated within the current time window to obtain rotation prediction, and the tilt angle is calibrated by gravity direction; the initial visual value is jointly modeled with the angular velocity and tilt angle constraints to minimize the visual reprojection error and electromechanical deviation term, and the rotation matrix R and translation vector T at the current moment are solved in real time; S45. The calculated R and T are subjected to time interpolation and noise suppression smoothing to maintain the consistency of the current timestamp and suppress instantaneous jitter, and the output is used as the spatial pose result; S46. At the same timestamp of the output R and T, the visual features at the corresponding moment are extracted from the visual data.

[0015] Optionally, the consistency verification method of S42 specifically includes:

[0016] (1) For each match, the most similar corresponding point of “frame t → frame t+1” and the most similar corresponding point of “frame t+1 → frame t” must be the same pair; if they are not the same pair, they are discarded.

[0017] (2) Based on the time interval and the known sampling frequency, set upper and lower limits for the pixel displacement of the matching point; remove those that exceed the reasonable range or have a sudden change in direction;

[0018] (3) Normalize the coordinates of the matching points in the two frames using the camera calibration parameters; estimate the fundamental matrix or essential matrix of the two frames under these coordinates, and calculate the distance from each pair of matches to the corresponding epipolar line; discard matches that are greater than the threshold.

[0019] (4) The iterative method of “random sampling - model estimation - error assessment - updating the interior point set” is adopted. The maximum interior point set that meets the epipolar geometric error threshold is retained in multiple rounds, and other matches are regarded as outliers and removed.

[0020] (5) Based on the fundamental matrix or essential matrix estimated in step (2), calculate the reprojection error for the set of matching points after filtering in steps (1) and (2); when the error exceeds the set pixel threshold, it is removed, and the remaining ones are the final set of feature points.

[0021] Optionally, S5 specifically includes:

[0022] S51. Within the time window corresponding to the dynamic attitude estimation module, extract the normalized rotation speed, vibration and attitude angle change parameters from the multi-source data sequence output by the time synchronization and data preprocessing module, and synchronously obtain the rotation matrix and translation vector corresponding to the time window, while extracting the visual features of the current frame.

[0023] S52. Based on the rotation matrix and translation vector, map the visual features of the current frame to the reference coordinate system to achieve spatial alignment of visual features under different pose conditions;

[0024] S53. Perform statistical analysis on the electromechanical signal characteristic parameters output by the time synchronization and data preprocessing module, calculate the attitude angle change rate, speed change rate and vibration amplitude change. When the attitude angle change rate or vibration amplitude exceeds the set threshold, it is determined to be a dynamic working condition. When the parameter changes are stable, it is determined to be a stable working condition. When it is between the two, it is determined to be a transitional working condition.

[0025] S54. Automatically set geometric constraint weights according to the operating conditions. When in a dynamic operating condition, the system sets the geometric constraint weights to the first weight value interval within the preset weight range, and the participation ratio of the spatial mapping constraint term in the feature optimization calculation increases to above the set ratio threshold. When in a stable operating condition, the system sets the geometric constraint weights to the second weight value interval within the preset weight range, so that the participation ratio of the spatial mapping constraint term in the feature optimization calculation is controlled below the set ratio threshold. When in a transitional operating condition, the system calculates the current constraint weights by time series interpolation between the first weight value interval and the second weight value interval.

[0026] Furthermore, the system records the previous weight output value within each time window and performs smooth transition calculations based on the current working condition identification results; when the weight change rate within a continuous window exceeds the set limit, the system automatically introduces a memory factor to limit the weight jump amplitude.

[0027] The adaptive introduction mechanism of the memory factor includes four states: enabled, enhanced, observed, and revoked. The system automatically switches between the four states based on the magnitude of weight candidate changes, visual stability, cross-modal synchronization quality, and electromechanical impact level. When in the enabled or enhanced state, the single-step change of weight is limited and directional consistency constraints are executed. When entering the observed state and the recovery criterion is met for multiple consecutive time windows, the memory factor is automatically revoked and the system reverts to the regular weight update strategy.

[0028] Optionally, the joint embedding fusion in step S6 specifically includes:

[0029] S61. Receive the visual features after dynamic geometric constraint consistency processing and the electromechanical signal features extracted based on the original electromechanical signals, and perform timestamp matching and dimension standardization processing on the two types of features.

[0030] S62. Perform feature mapping processing on visual features and electromechanical signal features respectively. Transform the visual features into visual embedding vectors in a unified feature space through a nonlinear mapping function, and transform the electromechanical signal features into electromechanical embedding vectors of the same dimension through a weighted linear mapping function.

[0031] S63. Perform temporal alignment and correlation modeling on the visual embedding vector and the electromechanical embedding vector, and obtain the feature coupling degree between the two by calculating the correlation within the same time window; allocate fusion coefficients according to the feature coupling degree.

[0032] S64. Jointly concatenate the visual embedding vector and the electromechanical embedding vector in the time-corresponding dimension, and introduce a fusion coefficient for weighted combination to obtain preliminary fusion features; perform feature normalization and principal component reduction on the preliminary fusion features.

[0033] S65. Perform sliding aggregation calculation on the reduced fusion features within the time window to extract the temporal correlation pattern reflecting the trend of winch state change; the sliding aggregation result is nonlinearly mapped to generate the final multimodal feature vector;

[0034] S66. Output the generated multimodal feature vector to the intelligent recognition module.

[0035] Optionally, the classification and regression analysis in step S7 specifically includes:

[0036] S71. Verify the consistency between the timestamp and the previous time; when missing or outlier values ​​occur, use interpolation and limiting methods within the same time window to fill in and constrain them.

[0037] S72. Within the current time window, calculate the mean, variation amplitude, and trend of the multimodal feature vector within the window to form a feature set for discrimination and localization at the current moment; suppress short-term jitter through time smoothing.

[0038] S73. Set up a classification subunit within the intelligent recognition module for distinguishing operating status and anomaly types:

[0039] (a) Calculate the discrimination score for each candidate category for the current feature set; (b) Generate category confidence based on the discrimination score and a preset threshold; (c) Output the corresponding running status and anomaly type when the confidence meets the threshold condition; If there are multiple categories that are close, the final category is determined by the weighted rule of the previous result and the current score.

[0040] S74. Set up a regression sub-unit for fault location and severity estimation in the intelligent recognition module: (a) Calculate the continuous output quantity related to the location (such as the relative position index along the roll width direction) for the current feature set; (b) Use residual constraints and time smoothing to suppress the influence of single frame error on the location output; (c) When the location output crosses the preset threshold interval, it is determined as a valid positioning result and stored in association with the classification result.

[0041] S75. Perform consistency checks on classification output and regression output. When the category indicates "normal" but the location output continues to exceed the limit, maintain the observation state and extend the time window. When the category is abnormal but the location output stably points to the same area, confirm the abnormality and output the location result. When the two conflict, the one with higher cumulative stability over time shall prevail and a check mark shall be recorded.

[0042] S76. Statistically analyze the discrimination and location results within multiple time windows of the nearest neighbors, and output the final category, fault location and corresponding confidence level at the current moment; when the confidence level is insufficient, maintain the previous reliable result and mark it as pending confirmation.

[0043] S77. Output the final operating status, anomaly type, fault location, and confidence level; when the result meets the early warning conditions, send the result and level information to the early warning and control module.

[0044] Optionally, S8 specifically includes:

[0045] S81. Calculate the comprehensive risk score based on the comprehensive identification result set and the stability index of electromechanical signal features and visual features within the corresponding time window;

[0046] S82. The comprehensive risk score is compared with the preset risk threshold, and divided into safe operation level, warning level, risk warning level and emergency warning level according to the score range;

[0047] S83. Output corresponding early warning signals according to the risk level.

[0048] The beneficial effects of this invention are:

[0049] (1) By introducing an adaptive feature modeling mechanism based on dynamic geometric consistency constraints, the present invention enables visual features to maintain spatial consistency and temporal stability under complex working conditions such as winch rotation, translation and vibration, which significantly improves the accuracy of posture recognition and feature extraction.

[0050] (2) This invention adopts a multimodal feature embedding and fusion method to fuse visual features and electromechanical signal features in a unified feature space, establishes the correspondence between visual information and electromechanical response, and realizes a comprehensive representation of the overall operating status of the winch. The system can automatically identify the working condition type according to the electromechanical signal features and adaptively adjust the geometric constraint weights so that the model can maintain optimal recognition performance under stable, transitional and dynamic working conditions.

[0051] (3) The present invention also combines a multi-level risk scoring and graded early warning mechanism, which can identify the operating status in real time and output the corresponding early warning level. In case of serious abnormality, it can automatically link and control the winch to perform deceleration or shutdown operations, thereby realizing integrated closed-loop management of monitoring and safety control. Attached Figure Description

[0052] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings,

[0053] Figure 1 This is a flowchart of a machine vision-based intelligent monitoring system for winches proposed in this invention.

[0054] Figure 2 This is a flowchart of the dynamic attitude estimation of a winch intelligent monitoring system based on machine vision proposed in this invention.

[0055] Figure 3 This is a flowchart of the adaptive feature modeling process for an intelligent monitoring system for winches based on machine vision, as proposed in this invention. Detailed Implementation

[0056] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0057] refer to Figures 1 to 3 A machine vision-based intelligent monitoring system for winches includes:

[0058] The visual acquisition module acquires continuous image sequences of the winch drum and wire rope during operation, generating visual data. The electromechanical signal acquisition module synchronously acquires raw electromechanical signals of the winch, including current, speed, tension, vibration, and attitude angles. The time synchronization and data preprocessing module performs time alignment, filtering, and normalization on the visual data and electromechanical signals, generating a multi-source data sequence. The dynamic attitude estimation module calculates the winch's rotation matrix and translation vector in real time based on changes in angular velocity, tilt angle, and visual feature points in the multi-source data sequence, representing its spatial attitude; it also extracts visual features from the visual data to provide input for subsequent modeling. Based on dynamic geometry... The adaptive feature modeling mechanism with geometric consistency constraints is used to maintain visual feature consistency under attitude change conditions. It performs spatial alignment based on rotation matrix and translation vector, and dynamically adjusts constraint weights according to the operating status. The multimodal feature fusion module fuses visual features processed by geometric consistency constraints with electromechanical signal features to generate multimodal feature vectors reflecting the operating status of the winch. The intelligent recognition module identifies the operating status, anomaly type, and fault location of the winch based on the multimodal feature vectors, forming a comprehensive recognition result set. The early warning and control module outputs risk warnings based on the comprehensive recognition result set and controls the winch to decelerate or stop in the event of a serious anomaly.

[0059] In this embodiment, the modules are interconnected using the following method:

[0060] S1. Acquire a continuous image sequence of the winch drum and wire rope during operation to generate visual data; S2. Synchronously acquire raw electromechanical signals of current, rotation speed, tension, vibration, and attitude angle; S3. Perform time alignment, filtering, and normalization on the visual data and raw electromechanical signals to generate a multi-source data sequence; S4. Based on the changes in angular velocity, tilt angle, and visual feature points of adjacent frames in the multi-source data sequence, calculate the rotation matrix and translation vector in real time; extract the visual features at the corresponding time from the visual data;

[0061] S5. An adaptive feature modeling mechanism based on dynamic geometric consistency constraints spatially aligns visual features according to the rotation matrix and translation vector, and dynamically adjusts the geometric constraint weights based on the electromechanical signal feature parameters of speed, vibration, current, and attitude angle provided by the electromechanical signal acquisition module and the time synchronization and data preprocessing module; S6. The visual features and electromechanical signal features are jointly embedded and fused to generate a multimodal feature vector reflecting the overall operating status of the winch; the electromechanical signal features are feature results extracted based on the original electromechanical signals, including mean, variance, frequency domain energy, or time-series rate of change; S7. Classification and regression analysis are performed based on the multimodal feature vector, and the operating status identification result, anomaly type identification result, and fault location identification result of the winch are output respectively, forming a comprehensive identification result set; S8. A risk classification early warning signal is output based on the comprehensive identification result set, and the winch is controlled to execute deceleration or shutdown commands in the event of a serious anomaly.

[0062] In this embodiment, S4 specifically includes:

[0063] S41. Select a data window for the current time and adjacent times from a multi-source data sequence according to a unified timestamp;

[0064] S42. Detect corner points or edge visual feature points on adjacent image frames respectively, and calculate the correspondence between adjacent frames based on similarity measurement to form a matching result set composed of visual feature points; use a consistency check method to remove matching points that do not meet geometric constraints to obtain a usable feature point correspondence set; the consistency check is as follows:

[0065] (1) For each match, the most similar corresponding point of “frame t → frame t+1” and the most similar corresponding point of “frame t+1 → frame t” must be the same pair; if they are not the same pair, they are discarded.

[0066] (2) Based on the time interval and the known sampling frequency, set upper or lower limits for the pixel displacement of the matching point; remove those that exceed the reasonable range or have a sudden change in direction;

[0067] (3) Normalize the coordinates of the matching points in the two frames using the camera calibration parameters; estimate the fundamental matrix or essential matrix of the two frames under these coordinates, and calculate the distance from each pair of matches to the corresponding epipolar line; discard matches that are greater than the threshold.

[0068] (4) The iterative method of “random sampling - model estimation - error assessment - updating the interior point set” is adopted. The maximum interior point set that meets the epipolar geometric error threshold is retained in multiple rounds, and other matches are regarded as outliers and removed.

[0069] (5) Based on the fundamental matrix or essential matrix estimated in step (2), calculate the reprojection error for the set of matching points after filtering in steps (1) and (2); when the error exceeds the set pixel threshold, it is discarded, and the remaining ones are the final set of feature points; S43, use the camera calibration parameters to normalize the feature points in the matching result set, and estimate the initial relative pose between adjacent image frames based on the normalized feature points to obtain the preliminary rotation parameters and translation direction parameters;

[0070] S43. Based on the device intrinsic parameters of the visual acquisition module, the feature points in the matching result set are normalized, and the initial relative pose values ​​between adjacent image frames are calculated based on the normalized feature points to obtain preliminary rotation parameters and translation direction parameters; S44. The angular velocity is numerically integrated within the current time window to obtain rotation prediction, and the tilt angle is calibrated by gravity direction; the initial visual value is jointly modeled with the angular velocity and tilt angle constraints to minimize the visual reprojection error and electromechanical deviation term, and the rotation matrix R and translation vector T at the current moment are solved in real time; S45. The calculated R and T are subjected to time interpolation and noise suppression smoothing to maintain the consistency of the current timestamp and suppress instantaneous jitter, and the output is used as the spatial pose result; S46. At the same timestamp of the output R and T, the visual features at the corresponding moment are extracted from the visual data.

[0071] In this embodiment, the adaptive feature modeling based on dynamic geometric consistency constraints in step S5 specifically includes:

[0072] S51. Within the time window corresponding to the dynamic attitude estimation module, extract the normalized rotation speed, vibration and attitude angle change parameters from the multi-source data sequence output by the time synchronization and data preprocessing module, and synchronously obtain the rotation matrix and translation vector corresponding to the time window, while extracting the visual features of the current frame.

[0073] S52. Based on the rotation matrix and translation vector, map the visual features of the current frame to the reference coordinate system to achieve spatial alignment of visual features under different pose conditions;

[0074] S53. Perform statistical analysis on the electromechanical signal characteristic parameters output by the time synchronization and data preprocessing module, calculate the attitude angle change rate, speed change rate and vibration amplitude change. When the attitude angle change rate or vibration amplitude exceeds the set threshold, it is determined to be a dynamic working condition. When the parameter changes are stable, it is determined to be a stable working condition. When it is between the two, it is determined to be a transitional working condition.

[0075] S54. Automatically set geometric constraint weights according to the operating conditions. When in a dynamic operating condition, the system sets the geometric constraint weights to the first weight value interval within the preset weight range, and the participation ratio of the spatial mapping constraint term in the feature optimization calculation increases to above the set ratio threshold. When in a stable operating condition, the system sets the geometric constraint weights to the second weight value interval within the preset weight range, so that the participation ratio of the spatial mapping constraint term in the feature optimization calculation is controlled below the set ratio threshold. When in a transitional operating condition, the system calculates the current constraint weights by time series interpolation between the first weight value interval and the second weight value interval.

[0076] In this embodiment, S5 specifically includes: the system records the previous weight output value in each time window and performs a smooth transition calculation based on the current working condition identification result; when the weight change rate in a continuous window exceeds a set limit, the system automatically introduces a memory factor to limit the weight jump amplitude.

[0077] The adaptive introduction mechanism of the memory factor includes four states: enabled, enhanced, observed, and revoked. The system automatically switches between the four states based on the magnitude of weight candidate changes, visual stability, cross-modal synchronization quality, and electromechanical impact level. When in the enabled or enhanced state, the single-step change of weight is limited and directional consistency constraints are executed. When entering the observed state and the recovery criterion is met for multiple consecutive time windows, the memory factor is automatically revoked and the system reverts to the regular weight update strategy.

[0078] In this embodiment, the joint embedding fusion in step S6 specifically includes:

[0079] S61. Receive the visual features after dynamic geometric constraint consistency processing and the electromechanical signal features extracted based on the original electromechanical signals, and perform timestamp matching and dimension standardization processing on the two types of features.

[0080] S62. Perform feature mapping processing on visual features and electromechanical signal features respectively. Transform the visual features into visual embedding vectors in a unified feature space through a nonlinear mapping function, and transform the electromechanical signal features into electromechanical embedding vectors of the same dimension through a weighted linear mapping function.

[0081] S63. Perform temporal alignment and correlation modeling on the visual embedding vector and the electromechanical embedding vector, and obtain the feature coupling degree between the two by calculating the correlation within the same time window; allocate fusion coefficients according to the feature coupling degree.

[0082] S64. Jointly concatenate the visual embedding vector and the electromechanical embedding vector in the time-corresponding dimension, and introduce a fusion coefficient for weighted combination to obtain preliminary fusion features; perform feature normalization and principal component reduction on the preliminary fusion features.

[0083] S65. Perform sliding aggregation calculation on the reduced fusion features within the time window to extract the temporal correlation pattern reflecting the trend of winch state change; the sliding aggregation result is nonlinearly mapped to generate the final multimodal feature vector;

[0084] S66. Output the generated multimodal feature vector to the intelligent recognition module.

[0085] In this embodiment, the classification and regression analysis in step S7 specifically includes:

[0086] S71. Verify the consistency between the timestamp and the previous time; when missing or outlier values ​​occur, use interpolation and limiting methods within the same time window to fill in and constrain them.

[0087] S72. Within the current time window, calculate the mean, variation amplitude, and trend of the multimodal feature vector within the window to form a feature set for discrimination and localization at the current moment; suppress short-term jitter through time smoothing.

[0088] S73. Set up a classification subunit within the intelligent recognition module for distinguishing operating status and anomaly types:

[0089] (a) Calculate the discrimination score for each candidate category for the current feature set; (b) Generate category confidence based on the discrimination score and a preset threshold; (c) Output the corresponding running status and anomaly type when the confidence meets the threshold condition; If there are multiple categories that are close, the final category is determined by the weighted rule of the previous result and the current score.

[0090] S74. Set up a regression sub-unit for fault location and severity estimation in the intelligent identification module: (a) Calculate the location-related continuous output for the current feature set;

[0091] (b) Residual constraints and time smoothing are used to suppress the impact of single-frame errors on the position output; (c) When the position output crosses the preset threshold range, it is determined to be a valid positioning result and stored in association with the classification result.

[0092] S75. Perform consistency checks on classification output and regression output. When the category indicates "normal" but the location output continues to exceed the limit, maintain the observation state and extend the time window. When the category is abnormal but the location output stably points to the same area, confirm the abnormality and output the location result. When the two conflict, the one with higher cumulative stability over time shall prevail and a check mark shall be recorded.

[0093] S76. Statistically analyze the discrimination and location results within multiple time windows of the nearest neighbors, and output the final category, fault location and corresponding confidence level at the current moment; when the confidence level is insufficient, maintain the previous reliable result and mark it as pending confirmation.

[0094] S77. Output the final operating status, anomaly type, fault location, and confidence level; when the result meets the early warning conditions, send the result and level information to the early warning and control module.

[0095] In this embodiment, S8 specifically includes:

[0096] S81. Calculate the comprehensive risk score based on the comprehensive identification result set and the stability index of electromechanical signal features and visual features within the corresponding time window;

[0097] S82. The comprehensive risk score is compared with the preset risk threshold, and divided into safe operation level, warning level, risk warning level and emergency warning level according to the score range;

[0098] S83. Output corresponding early warning signals according to the risk level.

[0099] Example 1

[0100] To verify the feasibility and superiority of this invention, it was applied to an intelligent monitoring system for winches designed to improve the safety of mechanical operation. The system is primarily used to automatically identify and intelligently warn of multi-dimensional parameters such as wire rope condition, drum speed, brake temperature, and load stress during winch operation, aiming to solve problems such as low efficiency, data lag, and inability to provide real-time diagnostics in traditional winch manual inspections.

[0101] The system consists of a high-resolution visual camera module, an infrared temperature sensing module, a load strain acquisition module, a control unit, and an intelligent algorithm processing platform. The camera module is mounted on the side of the winch's main shaft to acquire real-time video streams of wire rope winding, drum rotation, and brake disc operation. The visual module has a resolution of 2560×1440 pixels, a frame rate of 30 frames per second, and can capture drum angle changes within 0.1 seconds. The infrared temperature module has a temperature measurement range of −20℃ to 180℃ and a measurement accuracy of ±0.5℃; the strain module has a range of ±5000με and a sampling frequency of 500Hz, used to reflect the stress on the wire rope. During monitoring, the system first uses a visual recognition algorithm to extract edges and segment the drum region of the image, automatically calculating the drum rotation speed and the wire rope layer spacing. If a winding offset exceeding 2.5 mm or a drum jitter amplitude exceeding 3° is detected, the system immediately triggers a level one warning. Subsequently, a multimodal fusion model synchronously integrates visual information with strain and temperature data, generating a comprehensive health index through dynamic weight optimization. Multiple field tests demonstrate that the system maintains over 95% recognition accuracy even under complex lighting and vibration conditions. To verify the reliability and real-time performance of this invention, a comparative test was conducted between the traditional single-point sensing monitoring method and the system of this invention. The tests included drum speed recognition error, wire rope wear detection accuracy, temperature anomaly response time, load stability coefficient, and overall alarm delay. The test period was 6 hours of continuous operation, during which three working conditions—light load (30% rated load), medium load (60% rated load), and heavy load (90% rated load)—were randomly applied, resulting in approximately 180,000 data points. Under light load conditions, the system of this invention accurately identified changes in drum speed with an average error of only 0.12 revolutions per minute, while the traditional system's error reached 0.46 revolutions per minute. Under medium load operation, the wire rope surface wear recognition rate remained at 96.8%, an improvement of approximately 12 percentage points compared to the traditional method. Under heavy load conditions, the average response time for brake temperature rise monitoring was 1.8 seconds, nearly 50% shorter than the traditional temperature acquisition delay. In addition, this system uses an adaptive illumination compensation algorithm to automatically adjust exposure parameters within a 500 lx range of illumination variation, ensuring that the monitored image remains clear and the recognition error fluctuation does not exceed ±1.5%.

[0102] Table 1. Performance Comparison Results of Intelligent Monitoring Systems for Winches

[0103] Operating conditions Detection method Rotational speed recognition error (revolutions or minutes) Steel wire rope wear identification rate (%) Temperature rise response time (seconds) Alarm delay (seconds) Data stability coefficient (0~1) Average recognition accuracy (%) Light load Traditional monitoring 0.46 83.5 3.4 2.6 0.79 85.2 Light load This invention system 0.12 95.7 1.6 1.1 0.92 96.9 In-process Traditional monitoring 0.51 84.8 3.8 2.9 0.77 86.1 In-process This invention system 0.15 96.8 1.7 1.3 0.93 97.2 Heavy load Traditional monitoring 0.59 81.2 4.1 3.2 0.74 84.7 Heavy load This invention system 0.18 94.9 1.8 1.4 0.91 96.4 Light changes Traditional monitoring 0.63 78.4 4.5 3.5 0.72 82.6 Light changes This invention system 0.19 93.6 1.9 1.3 0.90 95.7 Vibration interference Traditional monitoring 0.57 80.1 4.2 3.0 0.76 84.1 Vibration interference This invention system 0.17 94.1 1.8 1.2 0.91 96.1

[0104] The data in the table show that the system of this invention exhibits significant performance advantages under various operating conditions. The overall speed recognition error is controlled within 0.2 revolutions per minute, a reduction of approximately 65% ​​compared to traditional monitoring methods, indicating that the visual feature recognition and angle estimation algorithms possess high accuracy. The wire rope wear recognition rate exceeds 93% under all operating conditions, reaching a maximum of 96.8% under medium-load conditions, effectively avoiding omissions caused by manual identification. The average temperature rise response time is only 1.75 seconds, about half that of traditional systems, indicating a significant improvement in response sensitivity after the fusion of the infrared temperature module and intelligent algorithm. The alarm delay is stable at around 1.3 seconds, enabling near real-time anomaly alerts, which is crucial for preventing brake failure. The data stability coefficient remains above 0.9, showing that the system maintains high reliability even under complex environments such as vibration and changes in lighting. The overall recognition accuracy remains around 96%, demonstrating the excellent performance of machine vision and multi-sensor fusion in winch condition monitoring.

[0105] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A machine vision-based intelligent monitoring system for winches, characterized in that, include: Visual acquisition module: used to acquire continuous image sequences of the winch drum and wire rope during operation and generate visual data; Electromechanical signal acquisition module: used to synchronously acquire the raw electromechanical signals of the winch, including current, speed, tension, vibration, and attitude angle; Time synchronization and data preprocessing module: used to time-align, filter, and normalize the visual data and electromechanical signals to generate a multi-source data sequence; Dynamic attitude estimation module: based on the changes in angular velocity, tilt angle, and visual feature points in the multi-source data sequence, calculates the winch's rotation matrix and translation vector in real time to characterize its spatial attitude; and simultaneously extracts visual features from the visual data; An adaptive feature modeling mechanism based on dynamic geometric consistency constraints is used to maintain visual feature consistency under pose change conditions. It performs spatial alignment based on rotation matrix and translation vector, and dynamically adjusts constraint weights according to the running state. Multimodal feature fusion module: Combines visual features processed by geometric consistency constraints with electromechanical signal features to generate a multimodal feature vector reflecting the operating status of the winch; Intelligent recognition module: Identifies the operating status, abnormality type and fault location of the winch based on the multimodal feature vector, forming a comprehensive recognition result set; Early warning and control module: Outputs risk warnings based on the comprehensive recognition result set, and controls the winch to decelerate or stop in case of serious abnormalities.

2. The intelligent monitoring system for winches based on machine vision according to claim 1, characterized in that, The modules are connected in the following way: S1. Acquire a continuous image sequence of the winch drum and wire rope during operation to generate visual data; S2. Synchronously acquire raw electromechanical signals of current, rotation speed, tension, vibration, and attitude angle; S3. Perform time alignment, filtering, and normalization on the visual data and raw electromechanical signals to generate a multi-source data sequence; S4. Based on the changes in angular velocity, tilt angle, and visual feature points of adjacent frames in the multi-source data sequence, calculate the rotation matrix and translation vector in real time; extract the visual features at the corresponding time from the visual data; S5. An adaptive feature modeling mechanism based on dynamic geometric consistency constraints spatially aligns visual features according to the rotation matrix and translation vector, and dynamically adjusts the geometric constraint weights based on electromechanical signal feature parameters of rotation speed, vibration, current, and attitude angle provided by the electromechanical signal acquisition module and the time synchronization and data preprocessing module; S6. The visual features and electromechanical signal features are jointly embedded and fused to generate a multimodal feature vector reflecting the overall operating status of the winch; the electromechanical signal features are feature results extracted based on the original electromechanical signals, including mean, variance, frequency domain energy, or time-series rate of change; S7. Classification and regression analysis are performed based on the multimodal feature vector, and the operating status identification results, anomaly type identification results, and fault location identification results of the winch are output respectively, forming a comprehensive identification result set; S8. Output risk classification early warning signals based on the comprehensive identification result set, and in the event of a serious abnormality, link the control of the winch to execute deceleration or stop commands.

3. The intelligent monitoring system for winches based on machine vision according to claim 1, characterized in that, The real-time calculation in step S4 specifically includes: S41. Select a data window for the current time and adjacent times from a multi-source data sequence according to a unified timestamp; S42. Detect corner points or edge visual feature points on adjacent image frames respectively, and calculate the correspondence between adjacent frames based on similarity measurement to form a matching result set composed of visual feature points; use consistency test method to remove matching points that do not meet geometric constraints to obtain a set of usable feature point correspondences. S43. Based on the device intrinsic parameters of the visual acquisition module, the feature points in the matching result set are normalized, and the initial relative pose values ​​between adjacent image frames are calculated based on the normalized feature points to obtain preliminary rotation parameters and translation direction parameters; S44. The angular velocity is numerically integrated within the current time window to obtain rotation prediction, and the tilt angle is calibrated by gravity direction; the initial visual value is jointly modeled with the angular velocity and tilt angle constraints to minimize the visual reprojection error and electromechanical deviation term, and the rotation matrix R and translation vector T at the current moment are solved in real time; S45. The calculated R and T are subjected to time interpolation and noise suppression smoothing to maintain the consistency of the current timestamp and suppress instantaneous jitter, and the output is used as the spatial pose result; S46. At the same timestamp of the output R and T, the visual features at the corresponding moment are extracted from the visual data.

4. The intelligent monitoring system for winches based on machine vision according to claim 3, characterized in that, The consistency check specifically includes: (1) For each match, the most similar corresponding point of "frame t → frame t+1" and the most similar corresponding point of "frame t+1 → frame t" must be the same pair; if they are not the same pair, they are discarded. (2) Based on the time interval and the known sampling frequency, set upper and lower limits for the pixel displacement of the matching point; remove those that exceed the reasonable range or have a sudden change in direction; (3) Normalize the coordinates of the matching points in the two frames using the camera calibration parameters; estimate the fundamental matrix or essential matrix of the two frames under these coordinates, and calculate the distance from each pair of matches to the corresponding epipolar line; discard matches that are greater than the threshold. (4) The iterative method of "random sampling - model estimation - error assessment - updating the interior point set" is adopted. The maximum interior point set that meets the epipolar geometric error threshold is retained in multiple rounds, and other matches are regarded as outliers and removed. (5) Based on the fundamental matrix or essential matrix estimated in step (2), calculate the reprojection error for the set of matching points after filtering in steps (1) and (2); when the error exceeds the set pixel threshold, it is removed, and the remaining ones are the final set of feature points.

5. The intelligent monitoring system for winches based on machine vision according to claim 1, characterized in that, The adaptive feature modeling based on dynamic geometric consistency constraints in step S5 specifically includes: S51. Within the time window corresponding to the dynamic attitude estimation module, extract the normalized rotation speed, vibration and attitude angle change parameters from the multi-source data sequence output by the time synchronization and data preprocessing module, and synchronously obtain the rotation matrix and translation vector corresponding to the time window, while extracting the visual features of the current frame. S52. Based on the rotation matrix and translation vector, map the visual features of the current frame to the reference coordinate system to achieve spatial alignment of visual features under different pose conditions; S53. Perform statistical analysis on the electromechanical signal characteristic parameters output by the time synchronization and data preprocessing module, calculate the attitude angle change rate, speed change rate and vibration amplitude change. When the attitude angle change rate or vibration amplitude exceeds the set threshold, it is determined to be a dynamic working condition. When the parameter changes are stable, it is determined to be a stable working condition. When it is between the two, it is determined to be a transitional working condition. S54. Automatically set geometric constraint weights according to the operating conditions. When in a dynamic operating condition, the system sets the geometric constraint weights to the first weight value interval within the preset weight range, and the participation ratio of the spatial mapping constraint term in the feature optimization calculation increases to above the set ratio threshold. When in a stable operating condition, the system sets the geometric constraint weights to the second weight value interval within the preset weight range, so that the participation ratio of the spatial mapping constraint term in the feature optimization calculation is controlled below the set ratio threshold. When in a transitional operating condition, the system calculates the current constraint weights by time series interpolation between the first weight value interval and the second weight value interval.

6. The intelligent monitoring system for winches based on machine vision according to claim 1, characterized in that, Specifically, S5 includes: the system records the previous weight output value in each time window and performs smooth transition calculation based on the current working condition identification result; when the weight change rate in a continuous window exceeds the set limit, the system automatically introduces a memory factor to limit the weight jump amplitude.

7. The intelligent monitoring system for winches based on machine vision according to claim 6, characterized in that, The adaptive introduction mechanism of the memory factor includes four states: enabled, enhanced, observed, and revoked. The system automatically switches between the four states based on the magnitude of weight candidate changes, visual stability, cross-modal synchronization quality, and electromechanical impact level. When in the enabled or enhanced state, the single-step change of weight is limited and directional consistency constraints are executed. When entering the observed state and the recovery criterion is met for multiple consecutive time windows, the memory factor is automatically revoked and the system reverts to the regular weight update strategy.

8. The intelligent monitoring system for winches based on machine vision according to claim 1, characterized in that, The joint embedding fusion in step S6 specifically includes: S61. Receive the visual features after dynamic geometric constraint consistency processing and the electromechanical signal features extracted based on the original electromechanical signals, and perform timestamp matching and dimension standardization processing on the two types of features. S62. Perform feature mapping processing on visual features and electromechanical signal features respectively. Transform the visual features into visual embedding vectors in a unified feature space through a nonlinear mapping function, and transform the electromechanical signal features into electromechanical embedding vectors of the same dimension through a weighted linear mapping function. S63. Perform temporal alignment and correlation modeling on the visual embedding vector and the electromechanical embedding vector, and obtain the feature coupling degree between the two by calculating the correlation within the same time window; allocate fusion coefficients according to the feature coupling degree. S64. Jointly concatenate the visual embedding vector and the electromechanical embedding vector in the time-corresponding dimension, and introduce a fusion coefficient for weighted combination to obtain preliminary fusion features; perform feature normalization and principal component reduction on the preliminary fusion features. S65. Perform sliding aggregation calculation on the reduced fusion features within the time window to extract the temporal correlation pattern reflecting the trend of winch state change; the sliding aggregation result is nonlinearly mapped to generate the final multimodal feature vector; S66. Output the generated multimodal feature vector to the intelligent recognition module.

9. The intelligent monitoring system for winches based on machine vision according to claim 1, characterized in that, The classification and regression analysis in step S7 specifically includes: S71. Verify the consistency between the timestamp and the previous time; when missing or outlier values ​​occur, use interpolation and limiting methods within the same time window to fill in and constrain them. S72. Within the current time window, calculate the mean, variation amplitude, and trend of the multimodal feature vector within the window to form a feature set for discrimination and localization at the current moment; suppress short-term jitter through time smoothing. S73. Set up a classification subunit within the intelligent recognition module for distinguishing operating status and anomaly types: (a) Calculate the discrimination score for each candidate category for the current feature set; (b) Generate category confidence based on the discrimination score and a preset threshold; (c) Output the corresponding running status and anomaly type when the confidence meets the threshold condition; If there are multiple categories that are close, the final category is determined by the weighted rule of the previous result and the current score. S74. Set up a regression sub-unit for fault location and severity estimation within the intelligent recognition module; a. Calculate the continuous output quantity related to the location for the current feature set; b. Use residual constraints and time smoothing to suppress the impact of single-frame errors on the location output; c. When the location output crosses the preset threshold range, it is determined as a valid positioning result and stored in association with the classification result. S75. Perform consistency checks on classification output and regression output. When the category indicates "normal" but the location output continues to exceed the limit, maintain the observation state and extend the time window. When the category is abnormal but the location output stably points to the same area, confirm the abnormality and output the location result. When the two conflict, the one with higher cumulative stability over time shall prevail and a check mark shall be recorded. S76. Statistically analyze the discrimination and location results within multiple time windows of the nearest neighbors, and output the final category, fault location and corresponding confidence level at the current moment; when the confidence level is insufficient, maintain the previous reliable result and mark it as pending confirmation. S77. Output the final operating status, anomaly type, fault location, and confidence level; when the result meets the early warning conditions, send the result and level information to the early warning and control module.

10. The intelligent monitoring system for winches based on machine vision according to claim 1, characterized in that, S8 specifically includes: S81. Calculate the comprehensive risk score based on the comprehensive identification result set and the stability index of electromechanical signal features and visual features within the corresponding time window; S82. The comprehensive risk score is compared with the preset risk threshold, and divided into safe operation level, warning level, risk warning level and emergency warning level according to the score range; S83. Output corresponding early warning signals according to the risk level.