A machine vision-based non-contact roller screen vibration detection method and system

CN122820580APending Publication Date: 2026-09-25HEFEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610918902.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而,现有视觉检测方案在滚轴筛应用中仍存在技术瓶颈:其一,现有方案大多依赖RGB图像进行分析,对同一画面内多个检测目标使用统一缩放比例,受各部件与相机间距不同的影响,振动幅值计算结果存在较大偏差;其二,特征点检测易受设备表面纹理、环境光照变化影响,易受环境噪声干扰产生虚假特征,振动频率与幅值计算精度不足;其三,未结合滚轴筛不同部件的振动特性设置差异化判定阈值,难以精准区分正常振动与异常状态,检测结果可靠性较差

Benefits of technology

本发明公开了一种基于机器视觉的非接触式滚轴筛振动检测方法及系统。本发明旨在解决传统接触式振动传感技术面临的维护成本高、现场部署受限等现实痛点,同时克服现有二维单目视觉检测技术在复杂且恶劣工业环境下适应性不足的技术瓶颈。对于具备多层结构、部件存在深度差别的滚轴筛设备,现有技术难以同步完成不同距离待测部件的振幅与频率的精确测量。本发明采用非接触式的多模态视觉信息融合与深度学习模型,将传统停留在二维像素层面的特征提取和运动跟踪扩展至三维物理空间,提高了检测系统的空间自适应能力与抗环境噪声能力。本发明可在多目标、多深度、高粉尘干扰的复杂工况下,实现设备振动的高精度实时监测,提升了设备故障检测的准确性与可靠性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820580A_ABST
    Figure CN122820580A_ABST
Patent Text Reader

Abstract

The application discloses a non-contact roller screen vibration detection method and system based on machine vision. First, the image of the detected part of the roller screen is collected to make a data set, and a YOLO target detection model weight is trained and deployed. The multi-modal image is collected by an RGB-D camera and stored in a buffer, and a CUDA acceleration model is called to efficiently extract the ROI region of the detected part from the RGB image. Then, the depth data is accurately aligned to the RGB coordinate system through the internal and external parameters of the camera, the weak texture feature points are extracted through grid uniform sampling, and the distance scaling factor of each feature point based on the depth information is obtained through distance gradient mapping depth information, and the detection and filtering of the feature points are completed. The filtered feature points are tracked by the LK pyramid optical flow method, and the stability is ensured by combining the jump interception and adaptive exponential smoothing strategy; and the more accurate amplitude is calculated through the scaling factor. Through the non-contact multi-modal visual information fusion and deep learning, the application realizes real-time and accurate vibration monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and industrial equipment condition monitoring technology, specifically to a non-contact roller screen vibration detection method and system based on machine vision. Background Technology

[0002] In the current era of deep integration between Industry 4.0 and intelligent manufacturing, the requirements for production continuity are becoming increasingly stringent in fields such as coal washing, power generation coal transportation, and building material aggregate processing. As a core piece of equipment for material grading, the operating status of roller screens directly determines the efficiency of the production line. Faults such as main shaft breakage or bearing jamming not only require emergency shutdowns for repairs but also cause economic losses to enterprises. Against this backdrop, equipment maintenance has gradually shifted from reactive repairs to predictive early warning systems. Vibration status, as a core indicator for fault diagnosis, has led the industry to place higher demands on the real-time performance and accuracy of detection technologies.

[0003] Currently, vibration detection mainly relies on traditional contact sensing technology, which involves installing accelerometers, vibration transmitters, and other devices on key parts of the equipment to collect vibration signals in real time and transmit them to the control system for analysis. While this method can monitor vibration at specific points, it has significant limitations in complex industrial scenarios: Firstly, the sensors need to be rigidly connected to the equipment, requiring machine shutdown during installation and involving complex on-site wiring; secondly, the equipment operating environment is characterized by high dust levels, high humidity, and severe vibration, which can easily cause sensor signal drift or hardware damage, resulting in high overall maintenance costs; thirdly, single-point sensing can only reflect local vibration conditions and cannot cover the overall scenario of multiple components working together in a roller screen, easily leading to missed or false detections and making comprehensive monitoring difficult.

[0004] With the development of machine vision and deep learning technologies, non-contact visual inspection is gradually becoming a new direction for equipment condition monitoring. This technology acquires equipment images through industrial cameras and extracts vibration features by combining image recognition and motion analysis algorithms, enabling non-contact, large-area monitoring coverage. However, existing visual inspection solutions still face technical bottlenecks in roller screen applications: First, most existing solutions rely on RGB images for analysis, using a uniform scaling ratio for multiple detection targets within the same image. Due to the varying distances between different components and the camera, the vibration amplitude calculation results show significant deviations. Second, feature point detection is easily affected by equipment surface texture and changes in ambient lighting, and is susceptible to environmental noise interference that generates false features, resulting in insufficient accuracy in vibration frequency and amplitude calculations. Third, the lack of differentiated judgment thresholds based on the vibration characteristics of different components of the roller screen makes it difficult to accurately distinguish between normal vibration and abnormal states, leading to poor reliability of the detection results.

[0005] Therefore, how to construct a vibration detection method that integrates multimodal visual information, adapts to complex industrial environments, and combines detection accuracy and comprehensiveness has become a technical challenge that urgently needs to be solved in the field of intelligent monitoring of industrial equipment. Summary of the Invention

[0006] To address the shortcomings of existing technologies, the present invention aims to provide a non-contact roller screen vibration detection method and system based on machine vision.

[0007] To achieve the above objectives, the technical solution of the present invention is as follows: A non-contact vibration detection method for roller screens based on machine vision, the method comprising the following steps: S1. Collect images of the area to be detected by the roller screen, construct a dataset and train a deep learning model, and deploy the model weights locally; S2. Use an RGB-D camera to synchronously acquire RGB images and depth maps in real time and cache them. Call the deployed model to complete inference, identify the target from the latest frame RGB image and extract the corresponding ROI region; S3. Based on the intrinsic and extrinsic parameters of the RGB-D camera, perform coordinate transformation on the depth map to align it with the RGB image coordinate system; extract feature points within the ROI region, filter out false feature points using the effective depth threshold to obtain effective feature points, and calculate the distance scaling factor based on the depth parameters of the effective feature points. S4. The effective feature points are tracked using the LK pyramid optical flow method, and the trajectory is optimized by jump interception and adaptive exponential smoothing. The displacement data is processed in combination with the distance scaling factor, and the vibration amplitude and vibration frequency are calculated by spectrum analysis, and background stationary points are removed. S5. Compare the calculated vibration amplitude and vibration frequency with the preset vibration threshold of the corresponding part to determine the operating status of the roller screen and output the detection results.

[0008] Further, step S1 specifically includes: S11. Use a camera to acquire images of the parts to be detected on the roller screen, annotate and preprocess the images, and construct a dataset; the annotation information includes the type and location information of the parts to be detected. S12. Divide the dataset into a training set, a validation set, and a test set in a ratio of 7:2:1; train a deep learning model using the training set to obtain the optimal weight file; and evaluate the model performance based on the test set. S13. Convert the format of the optimal weight file obtained from training, and then deploy it to the local system for identification of the parts to be detected by the roller screen.

[0009] Further, step S2 specifically includes: S21. Use an RGB-D camera to acquire synchronized RGB image and depth map data, and store the acquired data in a buffer. S22. Monitor the data retention time in the buffer in real time, clear data with a retention time of more than 35ms, and keep only the latest frame data in the buffer. S23. The YOLOv8 model is used as the object detection model. The model is deployed and inferred through the OpenCV DNN module, and CUDA is called to accelerate the inference. The ROI region corresponding to the part to be detected is extracted from the latest frame RGB image. This region is used as the associated target range of the depth data. S24. Determine whether the target area is identified in the image; if the target area is identified, output the corresponding ROI region and category information to the subsequent module; if the target area is not identified, return to step S23 to continue execution.

[0010] Further, step S3 specifically includes: S31. Read the pre-calibrated RGB-D camera intrinsic parameter matrix and extrinsic parameter matrix. The matrix includes a depth camera intrinsic parameter matrix, an RGB camera intrinsic parameter matrix, and an extrinsic parameter matrix used to transform the depth camera coordinate system to the RGB camera coordinate system. The depth camera intrinsic parameter matrix The calculation formula is:

[0011] in, This indicates the horizontal focal length of the depth camera. This indicates the focal length of the depth camera in the vertical direction. This represents the pixel coordinates of the optical center of the depth camera's imaging plane. Represents the pixel coordinates of the optical center of the depth camera's imaging plane; The RGB camera intrinsic parameter matrix The calculation formula is:

[0012] in, This indicates the horizontal focal length of the RGB camera. This indicates the vertical focal length of the RGB camera. Represents the optical center pixel of the imaging plane of an RGB camera x Axis coordinates Represents the optical center pixel of the imaging plane of an RGB camera y Axis coordinates; The extrinsic parameter matrix used to transform the depth camera coordinate system to the RGB camera coordinate system is: [R|T]; Where R is a 3D rotation matrix, representing the angular deflection relationship between the depth camera and the RGB camera, used to correct the rotational attitude deviation between the two cameras; T is a translation vector, representing the spatial position offset between the depth camera and the RGB camera. S33. After aligning the RGB features with the depth data, extract the initial feature point set in the ROI region of the feature detection area extracted by the YOLO model. ; The ROI region of the area to be detected is divided into: A uniform grid is used, and the gray standard deviation of the image within each small grid is calculated using the following formula. :

[0013] in, is the average grayscale value of the grid, and N is the total number of pixels in the grid.

[0014] If the standard deviation of grid gray level If the grid area contains weak device textures or edges, the GoodFeaturesToTrack algorithm is used to extract local optima, and a spatial distance rejection mechanism is set: if the Euclidean distance between the newly extracted feature point and the existing feature point is greater than the minimum spacing, then the excessively dense redundant points are removed. S34. Eliminate false feature points based on the depth information threshold to obtain valid feature points; S35. Based on the effective feature point depth, obtain the scaling factor based on the distance gradient; S36. Store the two-dimensional pixel coordinates of the selected valid feature points, the matching depth distance, the distance scaling factor sf, and the unique ID corresponding to each valid feature point in the feature point tracking data structure and output them to the downstream module.

[0015] Further, step S32 specifically includes: S321. Depth map pixel conversion to depth camera 3D coordinates: Traverse each pixel in the depth map. Read the depth value corresponding to the pixel. Combined with depth camera intrinsics The three-dimensional spatial coordinates of the pixel in the depth camera coordinate system can be calculated using the following formula. :

[0016] in, The horizontal pixel coordinates of the depth map pixels; The vertical pixel coordinates of the depth map pixels; The coordinates of the horizontal optical center pixel in the intrinsic parameters of the depth camera; The vertical optical center pixel coordinates in the depth camera's intrinsic parameters; The horizontal focal length of the depth camera; The vertical focal length of the depth camera; The x-axis spatial coordinates of a 3D point in the depth camera coordinate system; The y-axis spatial coordinates of a 3D point in the depth camera coordinate system; This is the depth measurement value corresponding to the current depth map pixel; S322, Using the calibration extrinsic parameter matrix [R|T] Transform to RGB camera coordinate system ; in, , The X-axis coordinates of a 3D point in the RGB camera coordinate system. The Y-axis coordinates of a 3D point in the RGB camera coordinate system. The depth distance of a 3D point along the Z-axis in the RGB camera coordinate system; S323, based on RGB camera intrinsic parameter matrix Will Pixel coordinates projected into the RGB camera coordinate system Generate depth maps corresponding to the pixels.

[0017] Further, step S35 specifically includes: S351. Based on the tangent geometric relationship, calculate the physical field of view size of the target plane at the current depth:

[0018]

[0019] in, The total horizontal physical width of the imaging area. The vertical total physical height of the imaging area. The depth distance of the feature points. For the camera's inherent horizontal field of view, This is the camera's inherent vertical field of view. S352, Discrete the physical field of view size uniformly to a resolution of [resolution missing]. In a digital pixel array, determine the horizontal scaling factor. and vertical scaling factor , horizontal scaling factor and vertical scaling factor The combination is the distance scaling factor sf for the effective feature point; The horizontal scaling factor The formula for calculating the horizontal physical length corresponding to a single pixel is:

[0020] The vertical scaling factor The vertical physical length corresponding to a single pixel is calculated using the following formula:

[0021] in, This represents the total horizontal physical width of the imaging region at the current depth. This represents the total horizontal pixel resolution of the image; This is the depth distance corresponding to the effective feature point, that is, the physical distance from the feature point to the RGB-D camera; This refers to the horizontal field of view of an RGB-D camera. This represents the total vertical physical height of the imaging area at the current depth. This represents the total vertical pixel resolution of the image. This refers to the vertical field of view of the RGB-D camera.

[0022] Further, step S4 specifically includes: S41. The LK pyramid optical flow method is used to track feature points, fixed parameters for optical flow tracking are configured, and pixel displacement velocity is calculated based on the optical flow constraint equation. S42. Based on the consistency of the optical flow residual and boundary constraints, the tracking jump caused by illumination and texture abrupt changes is intercepted by the pixel jump distance threshold, and the smoothing coefficient is dynamically adjusted according to the historical displacement fluctuation of the feature points to perform adaptive exponential smoothing on the effective feature point coordinates. S43. First, extract the displacement vector amplitude after weighted smoothing in step S42, multiply it by the corresponding distance scaling factor sf to obtain the true physical amplitude; then, set the amplitude protection period in the early stage of tracking to filter the values ​​during the device startup phase; finally, calculate the average value of the true physical amplitude of all effective feature points and use it as the overall vibration amplitude of a single frame in the current detection area. S44. Buffer the overall vibration amplitude of a single frame output by S43 frame by frame to construct a continuous multi-frame physical amplitude time sequence. Perform FFT spectrum analysis on the sequence to extract frequency domain vibration features. S45. The peak frequency index of the power spectrum is corrected by three-point parabolic interpolation to solve the accurate vibration frequency, and stationary feature points with amplitude and frequency below the threshold in multiple consecutive frames are removed, and the set of effective tracking feature points is output. S46. Extract the precise vibration frequency of each effective feature point in the effective tracking feature point set, and smooth the time sequence frequency by using a time window weighted average method. Finally, output the stable overall vibration amplitude and vibration main frequency of each detection part.

[0023] Further, step S42 specifically includes: S421. Based on the consistency between the optical flow residual and the boundary constraint verification trajectory, determine the tracking validity: when the optical flow state is normal, if the optical flow residual is less than 25 and the feature point coordinates do not exceed the image boundary, then the tracking of the feature point in the current frame is determined to be valid. S422. Calculate the pixel jump distance of feature points in adjacent frames and compare it with the displacement threshold preset based on the maximum rotation speed of the roller screen to intercept tracking jumps caused by changes in lighting and texture. set up The coordinates of the feature point at time step are The original coordinates at time t obtained by the LK optical flow method are: The instantaneous pixel jump distance is calculated using the following formula. :

[0024] Based on the maximum mechanical speed of the roller screen, a maximum reasonable pixel displacement threshold is preset. ,like If the optical flow has a matching jump, the abnormal displacement data is discarded. S423. Extract the historical displacement sequence within the feature point sliding window and calculate the sequence standard deviation. Dynamically calculate the smoothing coefficient using a piecewise threshold mapping method. ; S424. Use the exponential smoothing formula to perform weighted smoothing on the coordinates of effective feature points and output the feature time series coordinates. The exponential smoothing formula is as follows:

[0025] in, The coordinates of the feature points output after smoothing in the current frame t are the coordinates of the final stable trajectory in this frame; These are the coordinates of feature points that have already undergone smoothing in the previous frame t-1, representing historical stable trajectories; For adaptive smoothing coefficients, 0 < <1; The coordinates of the original valid feature points retained after the jump interception and filtering in frame t are obtained by direct calculation using the optical flow method. Step S44 specifically includes: S441. Perform DC component removal and Hanning window weighting on the original amplitude timing obtained from the buffer using the following formula:

[0026] in, To complete the preprocessing timing after removing DC and Hanning window weighting; This represents the amplitude corresponding to the nth sampling point in the original amplitude time series. This represents the average of the entire original time series. This represents the number of valid sampling points in the original time series. S442. Perform deep zero-padding on the preprocessed sequence output from S441, adding zero values ​​to the end of the sequence to extend the sequence length to the preset optimal discrete Fourier transform size. ; S443, Fill with zeros to... The time series of lengths is subjected to Discrete Fourier Transform (DFT) to solve for the frequency domain coefficients X(k) corresponding to each discrete frequency point. Then, the vibration power spectrum P(k) at each frequency point is calculated using the following formula:

[0027] S444. Set the target vibration frequency band to suit the working conditions of the roller screen, extract all power values ​​within the target frequency band, and calculate the median power. And retrieve the maximum power peak within the frequency band. The signal-to-noise ratio is verified using the following formula:

[0028] in, Signal-to-noise ratio (SNR) It is a very small constant; S445. A minimum signal-to-noise ratio (SNR) threshold is preset. If the calculated SNR reaches this threshold, the current vibration signal is determined to be valid, and valid power spectrum data is output. Step S45 specifically includes: S451. Based on the effective power spectrum data output in step S44, retrieve the index corresponding to the maximum power spectrum value. The power values ​​of the two adjacent frequency points to the left and right of the index are used for high-precision parabolic interpolation to calculate the peak offset of the frequency point. The calculation formula is:

[0029] in, , , Indexes Power values ​​corresponding to adjacent frequency points on the left, center, and right; This represents the frequency peak offset. S452. Combining the sampling frequency and the optimal transformation size, the high-precision vibration frequency is obtained using the following formula after correction. :

[0030] in, Image sampling frequency, The optimal discrete Fourier transform size; S453. Set a static feature point removal mechanism, record the state of each feature point. If the amplitude or frequency of the same feature point is less than the preset judgment threshold for multiple consecutive frames, then the feature point is determined to be a static background point without vibration, the feature point is removed from the tracking data structure, and the set of valid tracking feature points is output.

[0031] Further, step S5 specifically includes: S51. Read the vibration amplitude and frequency of each detected area, and construct a detection result data structure to store the detection data based on each ROI area and the identified roller screen detection part ID. S52. Match and retrieve the part ID in the test result data structure with the standard part ID in the pre-stored database, and determine the vibration status of the equipment based on the vibration threshold of each part. S53. If all detection data for the current part are within the safe threshold range, the device will output a normal detection result. S54. If the detection data is within the critical range of the safety threshold, the ROI area will be retested, and a more stringent retest threshold will be used for comparison. If the data after retesting is still within the critical range of the safety threshold, the detection anomaly result will be output, and the corresponding ROI area number and location information will be fed back. S55. If the detection data exceeds the safety threshold range of the corresponding part, the abnormal detection result will be output and the corresponding ROI area number will be fed back.

[0032] The present invention also includes a machine vision-based non-contact roller screen vibration detection system, comprising: The model training and deployment module is used to collect images of the parts to be detected by the roller screen, build a dataset and train a deep learning model, and deploy the model weights locally. The multimodal image acquisition and ROI extraction module is used to acquire RGB images and depth maps in real time and cache them using an RGB-D camera, call the deployed model to complete inference, identify targets from the latest frame of RGB images and extract the corresponding ROI regions; The effective feature point and scaling factor solving module is used to perform coordinate transformation on the depth map based on the intrinsic and extrinsic parameters of the RGB-D camera to align it with the RGB image coordinate system; extract feature points within the ROI region, filter out false feature points using the effective depth threshold to obtain effective feature points, and calculate the distance scaling factor based on the depth parameters of the effective feature points; The feature trajectory optimization and vibration parameter calculation module is used to track the trajectory of the effective feature points using the LK pyramid optical flow method, optimize the trajectory through jump interception and adaptive exponential smoothing; process the displacement data in combination with the distance scaling factor, calculate the vibration amplitude and vibration frequency through spectrum analysis, and remove background stationary points. The vibration state determination module is used to compare the calculated vibration amplitude and vibration frequency with the preset vibration threshold of the corresponding part to determine the operating state of the roller screen and output the detection results.

[0033] Compared with the prior art, the advantages of the present invention are: This invention discloses a non-contact vibration detection method and system for roller screens based on machine vision. The invention aims to address the practical pain points of traditional contact vibration sensing technology, such as high maintenance costs and limited on-site deployment, while overcoming the technical bottleneck of insufficient adaptability of existing two-dimensional monocular vision inspection technology in complex and harsh industrial environments. For roller screen equipment with multi-layered structures and components with varying depths, existing technologies struggle to simultaneously and accurately measure the amplitude and frequency of components at different distances. This invention employs a non-contact multimodal visual information fusion and deep learning model, extending traditional feature extraction and motion tracking from the two-dimensional pixel level to three-dimensional physical space, improving the spatial adaptability and environmental noise resistance of the detection system. This invention enables high-precision real-time monitoring of equipment vibration under complex working conditions involving multiple targets, multiple depths, and high dust interference, improving the accuracy and reliability of equipment fault detection. Attached Figure Description

[0034] Figure 1 This is a flowchart of the vibration detection method in this invention; Figure 2 This is a flowchart of the identification model deployment in this invention; Figure 3 This is a flowchart of the image acquisition process in this invention; Figure 4 This is a diagram showing the effect of extracting the ROI region based on the deep learning object detection algorithm in this invention; Figure 5 This is a flowchart of the vibration detection site identification and ROI region extraction process in this invention; Figure 6 This is a diagram showing the alignment effect of the RGB image and depth map in this invention; Figure 7 This is a diagram showing the high-confidence feature point extraction results in this invention. Figure 8 This is a flowchart of the scaling factor calculation and feature point extraction process in this invention; Figure 9 This is a flowchart of the vibration detection process in this invention; Figure 10 This is a flowchart of the vibration data detection process in this invention; Figure 11 The image shows the testing results of the roller screen under actual working conditions. Figure 12 This is a displacement diagram of the roller screen motor. Detailed Implementation

[0035] To provide a better understanding of the structural features and effects achieved by the present invention, a detailed description is provided below, accompanied by preferred embodiments and accompanying drawings: To address the shortcomings of existing technologies, this invention proposes a non-contact roller screen vibration detection method based on machine vision. This method integrates depth information and RGB image information. For scenarios involving multiple objects at varying distances during a single roller screen inspection, it utilizes the depth map to extract scaling factors in real time, enabling simultaneous measurement of targets at different depths within the same field of view, thus providing distance-adaptive detection. This invention expands the detection dimension from a two-dimensional plane to a three-dimensional space using depth information, effectively filtering out false feature points caused by lighting variations and dust occlusion in two-dimensional images, improving the reliability of feature extraction. Simultaneously, considering the complex textures of the surface equipment and the potential for tracking anomalies due to changes in ambient lighting, this invention introduces trajectory consistency checks and jump interception mechanisms to ensure the effectiveness and stability of feature point tracking using the LK optical flow method.

[0036] like Figure 1 The method shown is a non-contact vibration detection method for roller screens based on machine vision. The method includes the following steps: S1. Collect images of the areas to be detected by the roller screen, construct a dataset and train a deep learning model, and deploy the model weights locally.

[0037] S2. Use an RGB-D camera to synchronously acquire RGB images and depth maps in real time and cache them. Call the deployed model to complete inference, identify the target from the latest frame RGB image and extract the corresponding ROI region.

[0038] S3. Based on the intrinsic and extrinsic parameters of the RGB-D camera, perform coordinate transformation on the depth map to align it with the RGB image coordinate system; extract feature points within the ROI region, filter out false feature points using the effective depth threshold to obtain effective feature points, and calculate the distance scaling factor based on the depth parameters of the effective feature points. S4. The effective feature points are tracked using the LK pyramid optical flow method, and the trajectory is optimized by jump interception and adaptive exponential smoothing. The displacement data is processed in combination with the distance scaling factor, and the vibration amplitude and vibration frequency are calculated by spectrum analysis, and background stationary points are removed. S5. Compare the calculated vibration amplitude and vibration frequency with the preset vibration threshold of the corresponding part to determine the operating status of the roller screen and output the detection results.

[0039] As a further improvement to the above scheme, in step S1, an image of the part to be detected by the roller screen is collected, a dataset is constructed, and the model weights are trained and deployed to the local project as the basis for image recognition.

[0040] Step S1 specifically includes: S11. Use a camera to acquire images of the parts to be detected on the roller screen, label and preprocess the images, and construct a dataset; the labeling information includes the type and location information of the parts to be detected.

[0041] S12. Divide the dataset into a training set, a validation set, and a test set in a ratio of 7:2:1; train a deep learning model using the training set to obtain the optimal weight file; and evaluate the model performance based on the test set.

[0042] S13. Convert the format of the optimal weight file obtained from training, and then deploy it to the local system for identification of the parts to be detected by the roller screen.

[0043] Step S1, deploying the model, is fundamental to implementing the vibration detection method described in this invention. Only by accurately identifying the target region can subsequent detection work be carried out. Therefore, this embodiment deploys a deep learning target detection model in a local project. This model can efficiently infer the Region of Interest (ROI) of the area to be detected within the image, providing data support for subsequent feature point extraction and distance scaling factor calculation within the ROI.

[0044] This embodiment uses the existing open-source deep learning object detection model YOLOv8 to identify the parts to be detected on a roller screen. This model is a general object detection model in the field of machine vision; this invention only adapts it to the dataset, trains it, converts its format, and deploys it in engineering, without improving the model algorithm itself. The model is used to extract the ROI regions of the image, which serve as the basis for subsequent feature extraction and vibration analysis.

[0045] like Figure 2 As shown, the complete process of model deployment in this embodiment is as follows: First, use a camera to acquire images of the roller screen, label the parts to be detected and their corresponding position information, organize the image and label data, and create a dataset; then, divide the dataset into training set, validation set and test set in a ratio of 7:2:1, use the dataset to train the model and save the optimal weight file, and evaluate the model performance on the test set; finally, convert the model weight file to ONNX format and deploy it to a local C++ project to realize the automatic detection of target parts of the roller screen.

[0046] As a further improvement to the above scheme, in step S2, multimodal images are acquired in real time using an RGB-D camera and stored in a buffer. The deployed CUDA accelerated model is then invoked to identify the target (the region to be detected) from the latest frame of the RGB image and extract the ROI region corresponding to the region to be detected. The multimodal image includes an RGB image and a depth map.

[0047] Step S2 specifically includes: S21. Use an RGB-D camera to acquire synchronized RGB image and depth map data, and store the acquired data in a buffer.

[0048] S22. Monitor the data retention time in the buffer in real time, and clear data with a retention time exceeding 35ms to ensure that the buffer only retains the latest frame data. Extract valid RGB-D synchronization data from the buffer and output it to the subsequent processing module. The processing flow for this step is as follows: Figure 3 As shown.

[0049] S23. In this embodiment, the YOLOv8 model is used as the object detection model. The OpenCV DNN module is used for model deployment and inference, and CUDA is called to accelerate inference. The ROI region corresponding to the area to be detected is extracted from the latest frame RGB image. This region serves as the associated target range for the depth data. The detection effect is as follows: Figure 4 As shown. Figure 4 The on-site industrial images show three roller screen drive motors. The model outputs blue detection boxes marking each motor target, along with corresponding confidence levels of 0.93, 0.96, and 0.98. All confidence levels are at a high level. Figure 4 As can be seen, the deep learning object detection model deployed using this invention can accurately identify roller screen motor components under actual factory conditions, and stably output complete and undetected ROI regions for the equipment, providing a reliable target range for subsequent multimodal image registration and feature point extraction. Through the above reasoning, the model inference time is reduced from 100-200ms to approximately 25ms.

[0050] S24. Determine whether the target area (i.e., the area to be vibrated and detected by the roller screen) is identified in the image; if the target area is identified, output the corresponding ROI region and category information to the subsequent module; if the target area is not identified, return to step S23 to continue execution. The process of this step is as follows: Figure 5 As shown.

[0051] As a further improvement to the above scheme, step S3 performs coordinate registration between the RGB image and the depth map, filters effective feature points, and calculates the distance scaling factor for pixel-to-physical displacement transformation. First, based on the intrinsic and extrinsic parameters of the RGB-D camera, the depth map is transformed to align with the coordinate system of the RGB image. Then, a gridded uniform sampling method is used to extract weak texture feature points from the ROI region. Next, the depth values ​​of each feature point are extracted, and based on a set effective depth threshold, false feature points caused by complex lighting and flying dust in the two-dimensional image are filtered out to obtain effective feature points. Finally, the distance parameters of the effective feature points are obtained through the depth map, and the distance scaling factor sf of each effective feature point is calculated.

[0052] Step S3 specifically includes: S31. Read the pre-calibrated RGB-D camera intrinsic and extrinsic matrix, the matrix including the depth camera intrinsic matrix, the RGB camera intrinsic matrix and the extrinsic matrix used to transform the depth camera coordinate system to the RGB camera coordinate system.

[0053] Specifically, before running this embodiment, the dual-target calibration of the RGB-D camera is completed in advance, and the three sets of calibration parameters are stored locally. Step S31 reads all the calibration data for subsequent coordinate transformation calculations. The depth camera intrinsic parameter matrix is ​​used for calculating the conversion of depth image pixels to 3D spatial coordinates; the RGB camera intrinsic parameter matrix is ​​used for projecting 3D spatial points onto the RGB image pixel plane; the extrinsic parameter matrix includes a rotation matrix and a translation vector to achieve spatial coordinate transformation between the two camera coordinate systems.

[0054] Furthermore, the depth camera intrinsic parameter matrix The calculation formula is:

[0055] in, This indicates the horizontal focal length of the depth camera. This indicates the focal length of the depth camera in the vertical direction. This represents the pixel coordinates of the optical center of the depth camera's imaging plane. This represents the pixel coordinates of the optical center of the depth camera's imaging plane.

[0056] The RGB camera intrinsic parameter matrix The calculation formula is:

[0057] in, This indicates the horizontal focal length of the RGB camera. This indicates the vertical focal length of the RGB camera. Represents the optical center pixel of the imaging plane of an RGB camera x Axis coordinates Represents the optical center pixel of the imaging plane of an RGB camera y Axis coordinates.

[0058] The extrinsic parameter matrix used to transform the depth camera coordinate system to the RGB camera coordinate system is: [R|T] Where R is a 3D rotation matrix, representing the angular deflection relationship between the depth camera and the RGB camera, used to correct the rotational attitude deviation between the two cameras; T is a translation vector, representing the spatial position offset between the depth camera and the RGB camera.

[0059] Step S32: Based on the intrinsic and extrinsic parameter matrices of the RGB-D camera, all pixels of the depth map are mapped one by one to the RGB image coordinate system through coordinate transformation, completing the pixel-level alignment of the depth map and the RGB image, and outputting the registered depth map, providing a foundation for subsequent feature point matching depth data within the ROI region. The alignment effect is as follows. Figure 6 As shown. Figure 6 The image on the left is the RGB image after recognition by the YOLOv8 model. The ROI areas of the two roller screen motors are marked with blue detection boxes and the detection confidence is marked. The image on the right is the pseudo-color depth map after coordinate alignment. The color gradient represents the depth value of each pixel from the camera. The motor target area is completely matched with the outline of the RGB image.

[0060] Step S32 specifically includes: S321. Depth map pixel conversion to depth camera 3D coordinates: Traverse each pixel in the depth map. Read the depth value corresponding to the pixel. Combined with depth camera intrinsics The three-dimensional spatial coordinates of the pixel in the depth camera coordinate system can be calculated using the following formula. :

[0061] in, The horizontal pixel coordinates of the depth map pixels; The vertical pixel coordinates of the depth map pixels; The coordinates of the horizontal optical center pixel in the intrinsic parameters of the depth camera; The vertical optical center pixel coordinates in the depth camera's intrinsic parameters; The horizontal focal length of the depth camera; The vertical focal length of the depth camera; The x-axis spatial coordinates of a 3D point in the depth camera coordinate system; The y-axis spatial coordinates of a 3D point in the depth camera coordinate system; This is the depth measurement value corresponding to the current depth map pixel.

[0062] S322, Using the calibration extrinsic parameter matrix [R|T] Transform to RGB camera coordinate system ; in, , The X-axis coordinates of a 3D point in the RGB camera coordinate system. The Y-axis coordinates of a 3D point in the RGB camera coordinate system. This represents the depth distance of a 3D point along the Z-axis in the RGB camera coordinate system.

[0063] S323, based on RGB camera intrinsic parameter matrix Will Pixel coordinates projected into the RGB camera coordinate system This generates a depth map corresponding to each pixel. In this embodiment, the alignment effect between the RGB image and the depth map is as follows: Figure 6 As shown.

[0064] S33. Mesh-based weak texture feature point extraction: After aligning RGB features with depth data, this invention extracts an initial feature point set in the ROI region of the feature detection area extracted by the YOLO model. .

[0065] Based on the aligned RGB image output by S32, and combined with the ROI region identified by the YOLO model in step S2, a uniformly distributed initial feature point set is extracted. This addresses the issues of weak surface texture and redundant local clustering of feature points on the equipment. To resolve the problem of weak surface texture on the object to be inspected (roller screen), this embodiment divides the ROI region of the area to be inspected into... A uniform grid is used, and the gray standard deviation of the image within each small grid is calculated using the following formula. :

[0066] in, is the average grayscale value of the grid, and N is the total number of pixels in the grid.

[0067] When the standard deviation of grid gray level When a grid region is determined to contain weak device textures or edges, it is deemed valuable for feature point extraction. For grids that meet the criteria, the GoodFeaturesToTrack algorithm is used to extract local optima, and a spatial distance rejection mechanism is set: if the Euclidean distance between a newly extracted feature point and existing feature points is greater than the minimum spacing, excessively dense redundant points are removed to ensure that the detected feature points uniformly cover the ROI region.

[0068] In this embodiment, the extracted high-confidence feature points are as follows: Figure 7 As shown in the figure, the green markers are the sampled texture feature points. In this embodiment, the target ROI region is divided into a uniform grid, and weak texture feature points are extracted within the grid. Figure 7 The visible feature points are uniformly distributed across the surface of the motor, while a small number of sampling points are distributed in the surrounding background area. Subsequent steps will involve using a depth threshold to filter out false background feature points, retaining only valid feature points on the device surface for optical flow tracking. Figure 7 As can be seen, the gridded sampling strategy adopted in this embodiment can extract texture feature points uniformly and densely on the component to be detected, ensuring sufficient data for vibration tracking and providing a reliable basis for subsequent displacement and vibration amplitude calculation.

[0069] Within the ROI region of the roller screen output by the YOLO model, sub-regions containing weak textures are screened by meshed grayscale discrimination, and a uniformly distributed high-confidence initial feature point set is extracted to solve the problems of weak surface texture, redundant feature point aggregation, and easy tracking failure of the roller screen equipment. The initial feature point set is then output to step S34 for depth threshold filtering.

[0070] S34. Eliminate false feature points based on the depth information threshold to obtain valid feature points.

[0071] In this embodiment, the initial feature point coordinates extracted in step S33 and the alignment depth map output in step S32 are used as inputs to distinguish between device entities and aerial interference sources in three-dimensional space, filtering out false feature points. Specifically, each initial feature point is queried. Depth values ​​in the aligned depth map Pre-set the effective depth range corresponding to the physical entity of the roller screen equipment. If the feature point depth If the point is not within the effective depth range, it is determined to be a false feature point caused by dust or reflection and is removed. The remaining points attached to the surface of the roller screen are the effective feature points, and the coordinates of the effective feature points and the corresponding depth values ​​are output to S35.

[0072] Roller screen equipment has a fixed physical spatial span, corresponding to an effective depth range. The coal dust flying around the scene is usually suspended in the air near the camera, and its corresponding depth meets the requirements. Environmental reflections and depth distortion areas exceed the required depth. Neither of them has valid information about the physical surface of the roller screen and therefore does not belong to the roller screen equipment body. Step S34 expands the feature selection from two-dimensional images to three-dimensional space, which can ensure that all the valid feature points retained in the end are attached to the surface of the roller screen body.

[0073] S35. Based on the effective feature point depth, obtain the scaling factor based on the distance gradient. Using the feature point depth distance Z and the camera's horizontal field of view Vertical field of view Image resolution The scaling factor for pixel to real physical displacement is derived to eliminate measurement errors of vibration amplitude of templates at different depths. The derivation is based on the pinhole imaging model. In the derivation process, it is necessary to assume that the camera optical axis is orthogonal to the target plane of the rolling screen.

[0074] Step S35 specifically includes: S351. Based on the tangent geometric relationship, calculate the physical field of view size of the target plane at the current depth:

[0075]

[0076] in, The total horizontal physical width of the imaging area. The vertical total physical height of the imaging area. The depth distance of the feature points. For the camera's inherent horizontal field of view, This is the camera's inherent vertical field of view.

[0077] In vision-based non-contact vibration measurement, Converting pixel displacements in image space into real physical displacements relies on constructing an accurate distance scaling factor sf. Based on a pinhole imaging model, and assuming the camera's optical axis is orthogonal to the target plane of the roller screen, this method utilizes the feature point depth distance Z and the camera's horizontal field of view. Vertical field of view Calculate the absolute physical field of view of the target plane.

[0078] S352, Discrete the physical field of view size uniformly to a resolution of [resolution missing]. In a digital pixel array, determine the horizontal scaling factor. and vertical scaling factor , horizontal scaling factor and vertical scaling factor The combination is the distance scaling factor sf for the effective feature point; The horizontal scaling factor The formula for calculating the horizontal physical length corresponding to a single pixel is:

[0079] The vertical scaling factor The vertical physical length corresponding to a single pixel is calculated using the following formula:

[0080] in, This represents the total horizontal physical width of the imaging region at the current depth. This represents the total horizontal pixel resolution of the image; This is the depth distance corresponding to the effective feature point, that is, the physical distance from the feature point to the RGB-D camera; This refers to the horizontal field of view of an RGB-D camera. This represents the total vertical physical height of the imaging area at the current depth. This represents the total vertical pixel resolution of the image. This refers to the vertical field of view of the RGB-D camera.

[0081] S36. The two-dimensional pixel coordinates, matching depth distance, distance scaling factor sf, and unique ID corresponding to each valid feature point are uniformly stored in the feature point tracking data structure and output to the downstream module. In this embodiment, the feature point extraction and RGB-D image alignment process is as follows: Figure 8 As shown.

[0082] As a further improvement to the above scheme, based on the valid feature point data output in step S3, which includes coordinates, depth, distance scaling factor sf, and unique ID, this embodiment uses the LK pyramid optical flow method for feature point temporal tracking. To address the tracking divergence problem caused by complex surface textures and changes in ambient lighting, this embodiment introduces trajectory consistency verification, jump interception, and adaptive exponential smoothing mechanisms to stabilize feature motion trajectories. Then, based on the scaling factor, the unique pixel value is converted into the actual physical vibration amplitude. Sub-pixel-level frequency determination is achieved through high-precision FFT spectrum analysis and three-point parabolic interpolation, and a stationary point removal mechanism is set to filter invalid features. Finally, the stable vibration amplitude and dominant frequency of each part of the roller screen are output after time-window weighted smoothing. This invention, by combining the LK pyramid optical flow method, jump interception, and adaptive smoothing strategy, can improve the stability of feature point trajectories. By employing high-precision fast Fourier transform for spectrum analysis and stationary point removal, it can improve detection accuracy.

[0083] Step S4 specifically includes: S41. The LK pyramid optical flow method is used to track feature points. Fixed parameters for optical flow tracking are configured, and pixel displacement velocity is calculated based on the optical flow constraint equation.

[0084] Specifically, using the feature point pixel coordinates output in step S36 as the initial input for tracking, the optical flow tracking fixed parameters are configured: the tracking window size is... The pyramid has 3 layers, and the optical flow constraint equation is:

[0085] Where (u, v) represents the pixel displacement velocity; These represent the horizontal and vertical grayscale gradients of the image, respectively. This represents the rate of change of pixel grayscale over time.

[0086] S42. Based on the consistency of the optical flow residual and boundary constraints, the tracking jump caused by illumination and texture abrupt changes is intercepted by the pixel jump distance threshold, and the smoothing coefficient is dynamically adjusted according to the historical displacement fluctuation of the feature points to perform adaptive exponential smoothing on the effective feature point coordinates.

[0087] Step S42 specifically includes: S421. Based on the consistency between the optical flow residual and the boundary constraint verification trajectory, determine the tracking validity: When the optical flow state is normal, if the optical flow residual is less than 25 and the feature point coordinates do not exceed the image boundary, then the tracking of the feature point in the current frame is determined to be valid.

[0088] S422. Calculate the pixel jump distance of feature points in adjacent frames and compare it with the displacement threshold preset based on the maximum rotation speed of the roller screen to intercept tracking jumps caused by changes in lighting and texture.

[0089] To address issues such as tracking point drift caused by changes in local texture due to coal slime accumulation on equipment surfaces or sudden changes in local lighting, this invention establishes a jump interception mechanism. The displacement vector of feature points between two adjacent frames is calculated. If the displacement magnitude exceeds a set tolerance threshold, a jump is identified and intercepted, discarding the abnormal displacement.

[0090] Specifically, let The coordinates of the feature point at time step are The original coordinates at time t obtained by the LK optical flow method are: The instantaneous pixel jump distance is calculated using the following formula. : .

[0091] Based on the maximum mechanical speed of the roller screen, a maximum reasonable pixel displacement threshold is preset. ,like If an optical flow mismatch occurs, the abnormal displacement data is discarded. Optical flow methods may experience mismatches (mismatches) or mismatches due to texture blurring, triggering a mismatch interception.

[0092] S423. Extract the historical displacement sequence within the feature point sliding window and calculate the sequence standard deviation. Dynamically calculate the smoothing coefficient using a piecewise threshold mapping method. The greater the fluctuation, the higher the standard deviation and the corresponding smoothing coefficient. The larger the value of , the more effective it is in suppressing instantaneous jitter by increasing the weight of historical trajectories.

[0093] S424. In order to enhance the smoothing and noise reduction effect, the effective feature point coordinates are weighted and smoothed using the exponential smoothing formula to adaptively suppress tracking jitter and output continuous and stable feature time-series coordinates.

[0094] The exponential smoothing formula is as follows:

[0095] in, The coordinates of the feature points output after smoothing in the current frame t are the coordinates of the final stable trajectory in this frame; These are the coordinates of feature points that have already undergone smoothing in the previous frame t-1, representing historical stable trajectories; For adaptive smoothing coefficients, 0 < <1; The coordinates of the original valid feature points retained after the jump interception and filtering in frame t are obtained by direct calculation using the optical flow method. The weighting coefficients for the new tracking coordinates in the current frame. The larger the value, the smaller the weight of the original coordinates in the current frame, and the stronger the smoothing and suppression effect.

[0096] The exponential smoothing formula is used for exponential moving average filtering. It relies on dynamic adjustment by weighted fusion of historical smoothed trajectories and effective tracking coordinates in the current frame to achieve this. The adaptive balance preserves the true vibration and suppresses tracking jitter caused by light and dust, outputting low-noise continuous time-series coordinates to ensure the accuracy of subsequent amplitude and spectrum calculations.

[0097] S43. First, extract the displacement vector amplitude after weighted smoothing in step S42, multiply it by the corresponding distance scaling factor sf to obtain the true physical amplitude, so as to convert the pixel displacement into the true physical displacement of the roller screen surface; then, set the amplitude protection period in the initial stage of tracking to filter the values ​​during the device startup stage to prevent violent shaking during startup; finally, calculate the average value of the true physical amplitude of all effective feature points and use it as the single-frame overall vibration amplitude of the current detection area.

[0098] S44. Buffer the overall vibration amplitude of a single frame output by S43 frame by frame to construct a time series sequence of physical amplitudes for multiple consecutive frames. Perform FFT spectrum analysis on the sequence to extract frequency domain vibration features.

[0099] Step S44 specifically includes: S441. The original amplitude timing obtained from the buffer is subjected to DC component removal and Hanning window weighting using the following formula to eliminate static DC offset and suppress spectral leakage caused by finite-length truncated signals:

[0100] in, To complete the preprocessing timing after removing DC and Hanning window weighting; This represents the amplitude corresponding to the nth sampling point in the original amplitude time series. The mean of the entire original timing sequence represents the DC component of the signal. Achieve DC component removal; This represents the number of valid sampling points in the original time series.

[0101] S442. Perform deep zero-padding on the preprocessed sequence output from S441, adding zero values ​​to the end of the sequence to extend the sequence length to the preset optimal discrete Fourier transform size. This allows for the encryption of frequency domain sampling points, refinement of the spectral curve, and improvement of frequency domain resolution.

[0102] S443, Fill with zeros to... The time series of lengths is subjected to Discrete Fourier Transform (DFT) to solve for the frequency domain coefficients X(k) corresponding to each discrete frequency point. Then, the vibration power spectrum P(k) at each frequency point is calculated using the following formula:

[0103] S444. Set the target vibration frequency band to suit the working conditions of the roller screen, extract all power values ​​within the target frequency band, and calculate the median power. And retrieve the maximum power peak within the frequency band. The signal-to-noise ratio is verified using the following formula:

[0104] in, Signal-to-noise ratio (SNR) It is a very small constant used to avoid the denominator being 0.

[0105] S445. A minimum signal-to-noise ratio (SNR) threshold is preset. If the calculated SNR reaches this threshold, the current vibration signal is determined to be valid, and valid power spectrum data is output.

[0106] S45. The peak frequency index of the power spectrum is corrected by three-point parabolic interpolation to solve the accurate vibration frequency. Static feature points with amplitude and frequency below the threshold in multiple consecutive frames are removed, and effective tracking feature points that can reflect the vibration of the equipment are retained. The set of effective tracking feature points is output.

[0107] Step S45 specifically includes: S451. Based on the effective power spectrum data output in step S44, retrieve the index corresponding to the maximum power spectrum value. The power values ​​of the two adjacent frequency points to the left and right of the index are used for high-precision parabolic interpolation to calculate the peak offset of the frequency point. The calculation formula is:

[0108] in, , , Indexes Power values ​​corresponding to adjacent frequency points on the left, center, and right; This represents the frequency peak offset.

[0109] S452. Combining the sampling frequency and the optimal transformation size, the high-precision vibration frequency is obtained using the following formula after correction. :

[0110] in, Image sampling frequency, This is the optimal discrete Fourier transform size.

[0111] S453. A stationary feature point removal mechanism is set up to record the state of each feature point. If the amplitude or frequency of the same feature point is less than a preset judgment threshold for multiple consecutive frames, the feature point is determined to be a background stationary point without vibration, and the feature point is removed from the tracking data structure, outputting a set of valid tracking feature points. This design can avoid invalid stationary points interfering with the overall vibration detection effect and improve the accuracy of vibration detection.

[0112] S46. Extract the precise vibration frequency of each effective feature point in the effective tracking feature point set. Smooth the temporal frequency using a time window weighted average method to suppress instantaneous random fluctuations. Finally, output the stable overall vibration amplitude and dominant vibration frequency of each detected part. The processing flow of this step is as follows: Figure 9 As shown.

[0113] As a further improvement to the above scheme, S5, the vibration amplitude and frequency of each detection part output in step S46 are compared with the preset vibration threshold of the corresponding detection part, and the vibration detection result is output.

[0114] The processing procedure in step S5 is as follows: Figure 10 As shown, the detection performance of this algorithm under actual working conditions is as follows: Figure 11 As shown. Figure 11 The red rectangle in the image is used to mark the target ROI area of ​​the roller screen motor, and the green dots inside the rectangle are the effective tracking feature points after screening is completed. Figure 11 The top left corner shows the overall detection results: the overall vibration amplitude of the device is 0.70mm, the distance between the camera and the target is 1.14~1.45m; the center of the motor area is marked with the device's main frequency of 1.30Hz, which is corrected by three-point parabolic interpolation, and T1, T2, T3, and T0 are key tracking feature markers in the area. Figure 11 This embodiment intuitively demonstrates the real-time performance and visualization capabilities of the method described, enabling simultaneous target identification, continuous tracking of feature points, and real-time calculation of vibration amplitude and dominant frequency in a real-world factory setting.

[0115] Step S5 specifically includes: S51. Read the vibration amplitude and frequency of each detected area, and construct a detection result data structure to store the detection data based on each ROI area and the identified roller screen detection part ID. S52. Match and retrieve the part ID in the test result data structure with the standard part ID in the pre-stored database, and determine the vibration status of the equipment based on the vibration threshold of each part. S53. If all the detection data of the current part are within the safe threshold range, the normal detection result of the equipment will be output. The vibration and amplitude reference standards of the roller screen motor and reducer are shown in Table 1. Table 1 contains the vibration data of the roller screen obtained from the detection of all effective feature points. As can be seen from Table 1, the vibration amplitude and main frequency data of each group meet the vibration detection requirements of the equipment.

[0116] Table 1

[0117] Figure 12 The curve shows the displacement of a single effective feature point on the surface of the roller screen motor as a function of the number of image frames. The horizontal axis represents the number of image acquisition frames, and the vertical axis represents the vibration displacement of the feature point. Figure 12 The curve in the figure exhibits a standard periodic sinusoidal fluctuation with stable amplitude and no obvious distortion or jump noise. Based on this time-displacement curve, spectral analysis yields the main vibration frequency of the equipment to be approximately 1.57Hz, which is consistent with the rated output vibration frequency of the roller screen motor on site. Figure 12 It can be seen that the optical flow tracking method used in this invention can stably collect the time-series data of the periodic vibration displacement of the equipment, and the displacement curve has a high signal-to-noise ratio. The vibration main frequency obtained based on this time-series calculation is consistent with the actual working conditions of the equipment.

[0118] S54. If the detection data is within the critical range of the safety threshold, the ROI area will be retested, and a more stringent retest threshold will be used for comparison. If the data after retesting is still within the critical range of the safety threshold, the detection anomaly result will be output, and the corresponding ROI area number and location information will be fed back. S55. If the detection data exceeds the safety threshold range of the corresponding part, the abnormal detection result will be output and the corresponding ROI area number will be fed back.

[0119] After completing the above steps, the vibration amplitude, vibration frequency, and vibration detection results of each detected part of the vibrating object can be accurately obtained. If continuous vibration detection is required, steps S2 to S5 are repeated cyclically. This invention achieves real-time monitoring of the vibration state of the tested object, high-precision numerical solution, and vibration anomaly early warning by acquiring multimodal images in real time, extracting weak texture features through meshing, adaptive noise-resistant optical flow tracing, and combining high-precision spectral analysis with parabolic interpolation and comparison of vibration thresholds for each part.

[0120] The present invention also includes a machine vision-based non-contact roller screen vibration detection system, comprising: The model training and deployment module is used to collect images of the parts to be detected by the roller screen, build a dataset and train a deep learning model, and deploy the model weights locally. The multimodal image acquisition and ROI extraction module is used to acquire RGB images and depth maps in real time and cache them using an RGB-D camera, call the deployed model to complete inference, identify targets from the latest frame of RGB images and extract the corresponding ROI regions; The effective feature point and scaling factor solving module is used to perform coordinate transformation on the depth map based on the intrinsic and extrinsic parameters of the RGB-D camera to align it with the RGB image coordinate system; extract feature points within the ROI region, filter out false feature points using the effective depth threshold to obtain effective feature points, and calculate the distance scaling factor based on the depth parameters of the effective feature points; The feature trajectory optimization and vibration parameter calculation module is used to track the trajectory of the effective feature points using the LK pyramid optical flow method, optimize the trajectory through jump interception and adaptive exponential smoothing; process the displacement data in combination with the distance scaling factor, calculate the vibration amplitude and vibration frequency through spectrum analysis, and remove background stationary points. The vibration state determination module is used to compare the calculated vibration amplitude and vibration frequency with the preset vibration threshold of the corresponding part to determine the operating state of the roller screen and output the detection results.

[0121] In summary, this invention discloses a non-contact roller screen vibration detection method and system based on machine vision, aiming to solve the problems of high maintenance cost, susceptibility to interference, and limitations of single-point monitoring in traditional contact sensing technologies. The method first collects images of the roller screen area to be detected to create a dataset, trains and deploys weights for target detection models such as YOLO, laying the foundation for image recognition. An RGB-D camera is used to synchronously acquire multimodal images in real time and store them in a buffer to ensure the timeliness of the latest frame data. Simultaneously, a CUDA-accelerated model is used to efficiently extract the ROI region of the area to be detected from the RGB images. Subsequently, depth data is precisely aligned to the RGB coordinate system using camera intrinsic and extrinsic parameters. Weak texture feature points are extracted using gridded uniform sampling, and distance scaling factors based on depth information are obtained by mapping depth information to distance gradients, completing feature point detection and filtering. The filtered feature points are tracked using the LK pyramid optical flow method, combined with jump interception and adaptive exponential smoothing strategies to ensure stability; a more accurate amplitude is calculated using scaling factors. This invention achieves real-time vibration monitoring through non-contact multimodal visual information fusion and deep learning, effectively overcoming the problems of poor adaptability, low efficiency and insufficient calculation accuracy of traditional detection methods in complex environments, and significantly improving the reliability of industrial equipment status early warning.

[0122] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.

Claims

1. A non-contact vibration detection method for roller screens based on machine vision, characterized in that, The method includes the following steps: S1. Collect images of the area to be detected by the roller screen, construct a dataset and train a deep learning model, and deploy the model weights locally; S2. Use an RGB-D camera to synchronously acquire RGB images and depth maps in real time and cache them. Call the deployed model to complete inference, identify the target from the latest frame RGB image and extract the corresponding ROI region; S3. Based on the intrinsic and extrinsic parameters of the RGB-D camera, perform coordinate transformation on the depth map to align it with the RGB image coordinate system; extract feature points within the ROI region, filter out false feature points using the effective depth threshold to obtain effective feature points, and calculate the distance scaling factor based on the depth parameters of the effective feature points. S4. The effective feature points are tracked using the LK pyramid optical flow method, and the trajectory is optimized by jump interception and adaptive exponential smoothing. The displacement data is processed in combination with the distance scaling factor, and the vibration amplitude and vibration frequency are calculated by spectrum analysis, and background stationary points are removed. S5. Compare the calculated vibration amplitude and vibration frequency with the preset vibration threshold of the corresponding part to determine the operating status of the roller screen and output the detection results.

2. The non-contact roller screen vibration detection method based on machine vision according to claim 1, characterized in that, Step S1 specifically includes: S11. Use a camera to acquire images of the parts to be detected on the roller screen, annotate and preprocess the images, and construct a dataset; the annotation information includes the type and location information of the parts to be detected. S12. Divide the dataset into a training set, a validation set, and a test set in a ratio of 7:2:1; train a deep learning model using the training set to obtain the optimal weight file; and evaluate the model performance based on the test set. S13. Convert the format of the optimal weight file obtained from training, and then deploy it to the local system for identification of the parts to be detected by the roller screen.

3. The non-contact roller screen vibration detection method based on machine vision according to claim 2, characterized in that, Step S2 specifically includes: S21. Use an RGB-D camera to acquire synchronized RGB image and depth map data, and store the acquired data in a buffer. S22. Monitor the data retention time in the buffer in real time, clear data with a retention time of more than 35ms, and keep only the latest frame data in the buffer. S23. The YOLOv8 model is used as the object detection model. The model is deployed and inferred through the OpenCV DNN module, and CUDA is called to accelerate the inference. The ROI region corresponding to the part to be detected is extracted from the latest frame RGB image. This region is used as the associated target range of the depth data. S24. Determine whether the target area is identified in the image; if the target area is identified, output the corresponding ROI region and category information to the subsequent module; if the target area is not identified, return to step S23 to continue execution.

4. The non-contact roller screen vibration detection method based on machine vision according to claim 3, characterized in that, Step S3 specifically includes: S31. Read the pre-calibrated RGB-D camera intrinsic parameter matrix and extrinsic parameter matrix. The matrix includes a depth camera intrinsic parameter matrix, an RGB camera intrinsic parameter matrix, and an extrinsic parameter matrix used to transform the depth camera coordinate system to the RGB camera coordinate system. The depth camera intrinsic parameter matrix The calculation formula is: in, This indicates the horizontal focal length of the depth camera. This indicates the focal length of the depth camera in the vertical direction. This represents the pixel coordinates of the optical center of the depth camera's imaging plane. Represents the pixel coordinates of the optical center of the depth camera's imaging plane; The RGB camera intrinsic parameter matrix The calculation formula is: in, This indicates the horizontal focal length of the RGB camera. This indicates the vertical focal length of the RGB camera. Represents the optical center pixel of the imaging plane of an RGB camera x Axis coordinates Represents the optical center pixel of the imaging plane of an RGB camera y Axis coordinates; The extrinsic parameter matrix used to transform the depth camera coordinate system to the RGB camera coordinate system is: [R|T]; Where R is a 3D rotation matrix, representing the angular deflection relationship between the depth camera and the RGB camera, used to correct the rotational attitude deviation between the two cameras; T is a translation vector, representing the spatial position offset between the depth camera and the RGB camera. S33. After aligning the RGB features with the depth data, extract the initial feature point set in the ROI region of the feature detection area extracted by the YOLO model. ; The ROI region of the area to be detected is divided into: A uniform grid is used, and the gray standard deviation of the image within each small grid is calculated using the following formula. : in, is the average grayscale value of the grid, and N is the total number of pixels in the grid; If the standard deviation of grid gray level If the grid area contains weak device textures or edges, the GoodFeaturesToTrack algorithm is used to extract local optima, and a spatial distance rejection mechanism is set: if the Euclidean distance between the newly extracted feature point and the existing feature point is greater than the minimum spacing, then the excessively dense redundant points are removed. S34. Eliminate false feature points based on the depth information threshold to obtain valid feature points; S35. Based on the effective feature point depth, obtain the scaling factor based on the distance gradient; S36. Store the two-dimensional pixel coordinates of the selected valid feature points, the matching depth distance, the distance scaling factor sf, and the unique ID corresponding to each valid feature point in the feature point tracking data structure and output them to the downstream module.

5. The non-contact roller screen vibration detection method based on machine vision according to claim 4, characterized in that, Step S32 specifically includes: S321. Depth map pixel conversion to depth camera 3D coordinates: Traverse each pixel in the depth map. Read the depth value corresponding to the pixel. Combined with depth camera intrinsics The three-dimensional spatial coordinates of the pixel in the depth camera coordinate system can be calculated using the following formula. : in, The horizontal pixel coordinates of the depth map pixels; The vertical pixel coordinates of the depth map pixels; The coordinates of the horizontal optical center pixel in the intrinsic parameters of the depth camera; The vertical optical center pixel coordinates in the depth camera's intrinsic parameters; The horizontal focal length of the depth camera; The vertical focal length of the depth camera; The x-axis spatial coordinates of a 3D point in the depth camera coordinate system; The y-axis spatial coordinates of a 3D point in the depth camera coordinate system; This is the depth measurement value corresponding to the current depth map pixel; S322, Using the calibration extrinsic parameter matrix [R|T] Transform to RGB camera coordinate system ; in, , The X-axis coordinates of a 3D point in the RGB camera coordinate system. The Y-axis coordinates of a 3D point in the RGB camera coordinate system. The depth distance of a 3D point along the Z-axis in the RGB camera coordinate system; S323, based on RGB camera intrinsic parameter matrix Will Pixel coordinates projected into the RGB camera coordinate system Generate depth maps corresponding to the pixels.

6. The non-contact roller screen vibration detection method based on machine vision according to claim 5, characterized in that, Step S35 specifically includes: S351. Based on the tangent geometric relationship, calculate the physical field of view size of the target plane at the current depth: in, The total horizontal physical width of the imaging area. The vertical total physical height of the imaging area. The depth distance of the feature points. For the camera's inherent horizontal field of view, This is the camera's inherent vertical field of view. S352, Discrete the physical field of view size uniformly to a resolution of [resolution missing]. In a digital pixel array, determine the horizontal scaling factor. and vertical scaling factor , horizontal scaling factor and vertical scaling factor The combination is the distance scaling factor sf for the effective feature point; The horizontal scaling factor The formula for calculating the horizontal physical length corresponding to a single pixel is: The vertical scaling factor The vertical physical length corresponding to a single pixel is calculated using the following formula: in, This represents the total horizontal physical width of the imaging region at the current depth. This represents the total horizontal pixel resolution of the image; This is the depth distance corresponding to the effective feature point, that is, the physical distance from the feature point to the RGB-D camera; This refers to the horizontal field of view of an RGB-D camera. This represents the total vertical physical height of the imaging area at the current depth. This represents the total vertical pixel resolution of the image. This refers to the vertical field of view of the RGB-D camera.

7. The non-contact roller screen vibration detection method based on machine vision according to claim 6, characterized in that, Step S4 specifically includes: S41. The LK pyramid optical flow method is used to track feature points, fixed parameters for optical flow tracking are configured, and pixel displacement velocity is calculated based on the optical flow constraint equation. S42. Based on the consistency of the optical flow residual and boundary constraints, the tracking jump caused by illumination and texture abrupt changes is intercepted by the pixel jump distance threshold, and the smoothing coefficient is dynamically adjusted according to the historical displacement fluctuation of the feature points to perform adaptive exponential smoothing on the effective feature point coordinates. S43. First, extract the displacement vector amplitude after weighted smoothing in step S42, multiply it by the corresponding distance scaling factor sf to obtain the true physical amplitude; then, set the amplitude protection period in the early stage of tracking to filter the values ​​during the device startup phase; finally, calculate the average value of the true physical amplitude of all effective feature points and use it as the overall vibration amplitude of a single frame in the current detection area. S44. Buffer the overall vibration amplitude of a single frame output by S43 frame by frame to construct a continuous multi-frame physical amplitude time sequence. Perform FFT spectrum analysis on the sequence to extract frequency domain vibration features. S45. The peak frequency index of the power spectrum is corrected by three-point parabolic interpolation to solve the accurate vibration frequency, and stationary feature points with amplitude and frequency below the threshold in multiple consecutive frames are removed, and the set of effective tracking feature points is output. S46. Extract the precise vibration frequency of each effective feature point in the effective tracking feature point set, and smooth the time sequence frequency by using a time window weighted average method. Finally, output the stable overall vibration amplitude and vibration main frequency of each detection part.

8. The non-contact roller screen vibration detection method based on machine vision according to claim 7, characterized in that, Step S42 specifically includes: S421. Based on the consistency between the optical flow residual and the boundary constraint verification trajectory, determine the tracking validity: when the optical flow state is normal, if the optical flow residual is less than 25 and the feature point coordinates do not exceed the image boundary, then the tracking of the feature point in the current frame is determined to be valid. S422. Calculate the pixel jump distance of feature points in adjacent frames and compare it with the displacement threshold preset based on the maximum rotation speed of the roller screen to intercept tracking jumps caused by changes in lighting and texture. set up The coordinates of the feature point at time step are The original coordinates at time t obtained by the LK optical flow method are: The instantaneous pixel jump distance is calculated using the following formula. : Based on the maximum mechanical speed of the roller screen, a maximum reasonable pixel displacement threshold is preset. ,like If the optical flow has a matching jump, the abnormal displacement data is discarded. S423. Extract the historical displacement sequence within the feature point sliding window and calculate the sequence standard deviation. Dynamically calculate the smoothing coefficient using a piecewise threshold mapping method. ; S424. Use the exponential smoothing formula to perform weighted smoothing on the coordinates of effective feature points and output the feature time series coordinates. The exponential smoothing formula is as follows: in, The coordinates of the feature points output after smoothing in the current frame t are the coordinates of the final stable trajectory in this frame; These are the coordinates of feature points that have already undergone smoothing in the previous frame t-1, representing historical stable trajectories; For adaptive smoothing coefficients, 0 < <1; The coordinates of the original valid feature points retained after the jump interception and filtering in frame t are obtained by direct calculation using the optical flow method. Step S44 specifically includes: S441. Perform DC component removal and Hanning window weighting on the original amplitude timing obtained from the buffer using the following formula: in, To complete the preprocessing timing after removing DC and Hanning window weighting; This represents the amplitude corresponding to the nth sampling point in the original amplitude time series. This represents the average of the entire original time series. This represents the number of valid sampling points in the original time series. S442. Perform deep zero-padding on the preprocessed sequence output from S441, adding zero values ​​to the end of the sequence to extend the sequence length to the preset optimal discrete Fourier transform size. ; S443, Fill with zeros to... The time series of lengths is subjected to Discrete Fourier Transform (DFT) to solve for the frequency domain coefficients X(k) corresponding to each discrete frequency point. Then, the vibration power spectrum P(k) at each frequency point is calculated using the following formula: S444. Set the target vibration frequency band to suit the working conditions of the roller screen, extract all power values ​​within the target frequency band, and calculate the median power. And retrieve the maximum power peak within the frequency band. The signal-to-noise ratio is verified using the following formula: in, Signal-to-noise ratio (SNR) It is a very small constant; S445. A minimum signal-to-noise ratio (SNR) threshold is preset. If the calculated SNR reaches this threshold, the current vibration signal is determined to be valid, and valid power spectrum data is output. Step S45 specifically includes: S451. Based on the effective power spectrum data output in step S44, retrieve the index corresponding to the maximum power spectrum value. The power values ​​of the two adjacent frequency points to the left and right of the index are used for high-precision parabolic interpolation to calculate the peak offset of the frequency point. The calculation formula is: in, , , Indexes Power values ​​corresponding to adjacent frequency points on the left, center, and right; This represents the frequency peak offset. S452. Combining the sampling frequency and the optimal transformation size, the high-precision vibration frequency is obtained using the following formula after correction. : in, Image sampling frequency, The optimal discrete Fourier transform size; S453. Set a static feature point removal mechanism, record the state of each feature point. If the amplitude or frequency of the same feature point is less than the preset judgment threshold for multiple consecutive frames, then the feature point is determined to be a static background point without vibration, the feature point is removed from the tracking data structure, and the set of valid tracking feature points is output.

9. The non-contact roller screen vibration detection method based on machine vision according to claim 1, characterized in that, Step S5 specifically includes: S51. Read the vibration amplitude and frequency of each detected area, and construct a detection result data structure to store the detection data based on each ROI area and the identified roller screen detection part ID. S52. Match and retrieve the part ID in the test result data structure with the standard part ID in the pre-stored database, and determine the vibration status of the equipment based on the vibration threshold of each part. S53. If all detection data for the current part are within the safe threshold range, the device will output a normal detection result. S54. If the detection data is within the critical range of the safety threshold, the ROI area will be retested, and a more stringent retest threshold will be used for comparison. If the data after retesting is still within the critical range of the safety threshold, the detection anomaly result will be output, and the corresponding ROI area number and location information will be fed back. S55. If the detection data exceeds the safety threshold range of the corresponding part, the abnormal detection result will be output and the corresponding ROI area number will be fed back.

10. A non-contact roller screen vibration detection system based on machine vision, characterized in that, include: The model training and deployment module is used to collect images of the parts to be detected by the roller screen, build a dataset and train a deep learning model, and deploy the model weights locally. The multimodal image acquisition and ROI extraction module is used to acquire RGB images and depth maps in real time and cache them using an RGB-D camera, call the deployed model to complete inference, identify targets from the latest frame of RGB images and extract the corresponding ROI regions; The effective feature point and scaling factor solving module is used to perform coordinate transformation on the depth map based on the intrinsic and extrinsic parameters of the RGB-D camera to align it with the RGB image coordinate system; extract feature points within the ROI region, filter out false feature points using the effective depth threshold to obtain effective feature points, and calculate the distance scaling factor based on the depth parameters of the effective feature points; The feature trajectory optimization and vibration parameter calculation module is used to track the trajectory of the effective feature points using the LK pyramid optical flow method, optimize the trajectory through jump interception and adaptive exponential smoothing; process the displacement data in combination with the distance scaling factor, calculate the vibration amplitude and vibration frequency through spectrum analysis, and remove background stationary points. The vibration state determination module is used to compare the calculated vibration amplitude and vibration frequency with the preset vibration threshold of the corresponding part to determine the operating state of the roller screen and output the detection results.