Hardware fitting size detection system and method based on visual detection

By constructing a visual inspection system that incorporates vibration sensing and spectrum decoupling, it can adaptively respond to complex disturbance scenarios, solving the problems of insufficient detection accuracy and stability in existing technologies, and achieving efficient and accurate detection of the dimensions of hardware accessories.

CN121230623BActive Publication Date: 2026-04-24HANGZHOU YIJIA INTELLIGENT HARDWARE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU YIJIA INTELLIGENT HARDWARE CO LTD
Filing Date
2025-11-28
Publication Date
2026-04-24

Smart Images

  • Figure CN121230623B_ABST
    Figure CN121230623B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of machine vision, and more particularly to a hardware fitting size detection system and method based on visual detection, the detection system comprising a vibration sensing module, an image acquisition module, a data processing and control module and a storage module, real-time acquisition of support structure vibration signals, matching with pre-stored standard scene models through spectrum analysis, identification of disturbance type; adaptive selection of wavelet decomposition strategy according to scene type, decoupling of vibration and impact components and enhancement of parameters, and then branched calculation of compensation exposure parameters; under complex disturbance, fusion of multi-weight strategy and feedback image closed-loop optimization of exposure time, and calling of scene-adapted image processing algorithm for size detection, while performing scene-bound rollback mechanism to restore features. The present application can solve the problem of insufficient detection accuracy and stability caused by the inability of a single processing model to cope with continuous vibration and transient impact coupling disturbance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of machine vision inspection technology, and in particular to a system and method for inspecting the dimensions of hardware accessories based on vision inspection. Background Technology

[0002] In the large-scale production process of hardware accessories, automated quality inspection has become a key link in ensuring product consistency and improving production efficiency. Machine vision-based online inspection technology, with its non-contact, high-speed, and high-precision characteristics, is widely used for real-time monitoring of key indicators such as workpiece dimensions and surface defects. However, with the continuous increase in production line speed and the increasing complexity of processes, the physical environment in which vision inspection systems operate is becoming increasingly demanding. Especially in high-speed conveyor belt systems, the continuous mechanical vibrations generated by the equipment itself, as well as the instantaneous impacts caused by conveyor belt seams or workpiece handover points, together constitute a complex, multi-frequency coupled disturbance scenario. This composite disturbance poses a severe challenge to the stability of image acquisition, directly affecting the accuracy and reliability of subsequent dimensional measurements.

[0003] To address these challenges, those skilled in the art have conducted numerous explorations and proposed corresponding solutions. For example, Chinese patent document CN118840312B discloses a machine vision-based method and system for detecting metal stamping defects. This solution acquires real-time images of the target part using a camera device and employs image enhancement processing and feature extraction algorithms to identify the stamping effect. Its design focuses on addressing the motion blur problem caused by single, continuous vibrations. By optimizing exposure parameters and using specific filters to sharpen the image, it suppresses the image quality degradation caused by steady-state equipment operation to a certain extent, which is significant for improving detection accuracy under single low-frequency vibration environments. Another Chinese patent, CN118552524B, discloses a visual detection method and system for metal stamping defects. Its technical effect is to perform detailed analysis of pixels at the edges of the image to detect abrupt changes in grayscale caused by defects such as burrs on the metal part, thereby achieving high-precision defect localization. These solutions have good capture capabilities for local features under static or micro-vibration environments, belonging to the approach of improving stability from the pixel-level signal level. The inherent characteristics of the aforementioned technical solutions in principle have gradually revealed their shortcomings when addressing new technical demands. All of these solutions are based on static or single-fluctuation models in their image processing and analysis logic. When faced with real-world industrial scenarios involving combined continuous vibration and transient impact excitation, these solutions reveal a profound fundamental technical contradiction: the image degradation caused by these different physical factors has inherently conflicting and incompatible underlying compensation mechanisms. Continuous vibration (conveyor belt vibration is typically 5-20Hz) causes motion blur in images, characterized by the blurring and diffusion of macroscopic features of target edges at the moment of imaging due to time integration. The core approach to handling this form of degradation is based on the continuity and compensability of motion, freezing the dynamic moment by shortening exposure time or using motion compensation. High-frequency transient impacts (greater than 100Hz) caused by conveyor belt seams, etc., cause positional jumps in the image, characterized by macroscopic and unpredictable spatial displacement of the target at the moment of imaging. The root cause of this form of degradation is sudden disturbances, and the essence of its compensation mechanism is the prediction and temporal synchronization of sudden disturbance events, such as avoiding displacement by triggering camera exposure in advance.

[0004] The aforementioned performance bottleneck arises because the two different perturbations acting on images under high-speed and complex conditions differ fundamentally in their physical parameters (frequency, amplitude, time) and the mechanisms affecting the imaging system (temporal domain integral blur vs. spatial domain instantaneous translation), as well as their corresponding optimization strategies. Existing methods simply use a fixed procedure to handle all perturbations, leading to a trade-off in performance. A system optimized to suppress motion blur cannot correct macroscopic positional shifts when encountering transient impacts, resulting in erroneous edge detection that further amplifies the error. Conversely, a system optimized for timing control cannot solve the persistent motion blur problem when faced with continuous vibrations. In complex scenarios, these two types of image degradation occur simultaneously and interact, making any simple compensation model ineffective. This leads to image feature destruction, causing a sudden decrease in detection accuracy and an increase in false detection rate. The paradigm based on static or single perturbation models has become the fundamental obstacle preventing further improvement in visual inspection technology under high-speed and complex conditions.

[0005] Therefore, how to design a system that can dynamically sense the complex disturbance environment of the production site, accurately decouple the different effects of continuous vibration and transient impact on the imaging process, and adaptively call the optimal image acquisition and processing strategy to fundamentally solve the problem of insufficient detection accuracy and stability caused by the single processing model in existing technologies has become a key challenge and an urgent technical problem for those skilled in the art. Summary of the Invention

[0006] The purpose of this invention is to provide a vision-based system and method for detecting the dimensions of hardware accessories, which can solve or at least alleviate the problem of insufficient detection accuracy and stability caused by the inability to effectively cope with the combined disturbance scenarios of continuous vibration and transient impact due to the use of a single, fixed processing model.

[0007] To achieve the above objectives, the present invention provides the following technical solution: a vision-based hardware accessory dimension inspection system, the system comprising:

[0008] The vibration sensing module is used to collect vibration signals from the system's support structure in real time.

[0009] The image acquisition module is used to acquire images of hardware accessories;

[0010] The storage module has a historical scene database and a set of compensation rules pre-stored inside it. The historical scene database defines at least two spectral feature models corresponding to standard disturbance scenarios of continuous steady-state vibration and transient impact, respectively.

[0011] A data processing and control module, which is electrically connected to the vibration sensing module, the image acquisition module, and the storage module, is configured as follows:

[0012] Receive vibration signals and perform spectrum analysis to generate real-time vibration spectrum;

[0013] The real-time vibration spectrum is matched with the spectral feature models of various standard disturbance scenarios in the historical scenario database to identify the type of disturbance scenario to which the current working condition belongs.

[0014] Based on the identified disturbance scenario type, load the corresponding compensation rule set;

[0015] Based on the loaded compensation rule set, image acquisition control parameters are generated for the current disturbance scene type; and

[0016] Based on the image acquisition control parameters, the image acquisition module is controlled to perform image acquisition, and according to the scene identifier in the image metadata, the corresponding scene-adaptive image processing algorithm is called to process the acquired image to calculate its size.

[0017] To further realize the present invention, the following technical solutions may be preferred:

[0018] Preferably, the vibration sensing module includes a triaxial microelectromechanical system accelerometer and an analog-to-digital converter. The accelerometer is mounted on the support structure of the image acquisition module, and the analog-to-digital converter converts the analog voltage signal into a digital vibration signal sequence.

[0019] Preferably, the historical scene database defines three standard disturbance scenarios:

[0020] The first standard scenario corresponds to continuous steady-state vibration, and its spectral characteristic model is defined as the presence of a dominant frequency peak in the low-frequency range.

[0021] The second standard scenario, corresponding to transient impact, is defined by its spectral characteristic model as the presence of a broadband pulse peak in the high-frequency range.

[0022] The third standard scenario, corresponding to the composite disturbance, has its spectral characteristic model defined as the spectral characteristics of both the first and second standard scenarios.

[0023] The compensation rule set stored in the storage module contains compensation rule sets that correspond one-to-one with the three standard disturbance scenarios.

[0024] Preferably, the data processing and control module is further configured to perform a scene-adaptive spectrum decoupling operation after identifying the disturbance scene type, the operation including:

[0025] If the current scene is identified as the first standard scene, the low-scale wavelet decomposition strategy is invoked to decompose the digital vibration signal sequence and reconstruct the continuous vibration components.

[0026] If the current scene is identified as the second standard scene, the high-scale wavelet decomposition strategy is invoked to decompose and reconstruct the transient impact component.

[0027] If the current scene is identified as the third standard scene, then the full-scale decomposition strategy is executed to reconstruct the continuous vibration component and the transient impact component.

[0028] Furthermore, the data processing and control module is further configured to perform scene enhancement calculations after validating the reconstructed components.

[0029] Preferably, the data processing and control module is configured to perform scene-bound branching processing operations to generate image acquisition control parameters, the operations including:

[0030] If the current scene is the first standard scene, then based on the enhanced continuous vibration component, the lookup table in the first compensation rule set is queried to obtain the basic exposure time, and the basic exposure time is adjusted according to the case where the amplitude exceeds the preset threshold.

[0031] If the current scene is the second standard scene, the enhanced transient impact component will be matched with the impact feature library in the second compensation rule set, the timestamp of the impact event will be recorded, and the early exposure trigger time point will be calculated.

[0032] If the current scene is the third standard scene, the processing logic of the first and second standard scenes above will be executed in parallel to generate the continuous vibration compensation exposure time and the impact compensation early trigger time point, respectively, and the parameter fusion operation will be performed to generate the final image acquisition control parameters.

[0033] A vision-based method for detecting the dimensions of hardware accessories, applicable to the aforementioned system, comprising the following steps:

[0034] The vibration signal of the system support structure is collected in real time by the vibration sensing module and converted into a digital vibration signal sequence.

[0035] Spectral analysis is performed on the digital vibration signal sequence, and the obtained real-time vibration spectrum is matched with the standard scene spectral feature model in the pre-stored historical scene database. This identifies the current disturbance environment as one of the first standard scene, the second standard scene, or the third standard scene, and loads the compensation rule set corresponding to the identified scene.

[0036] Based on the identified scene type, an appropriate wavelet transform decomposition strategy is selected and executed to decouple the continuous vibration component and / or transient impact component from the digital vibration signal sequence, and the parameters of the decoupled components are verified and scene enhancement calculations are performed.

[0037] Depending on the scene type, different processing branches are entered to calculate image acquisition control parameters;

[0038] Based on the calculated image acquisition control parameters, the image acquisition module is controlled to acquire images of hardware accessories, and scene identifiers are added to the image metadata;

[0039] Based on the scene identifier in the image metadata, the corresponding image processing algorithm is invoked to perform feature extraction and size calculation, and the results are output.

[0040] Preferably, in the step of calculating image acquisition control parameters, when the current disturbance environment is identified as a third standard scene, the method further includes a scene-driven parameter fusion operation, the parameter fusion operation including:

[0041] Using preset continuous vibration compensation weights and shock compensation weights, the exposure parameters implied by the parallel calculated continuous vibration compensation exposure time and the shock compensation early trigger time are weighted and fused.

[0042] The weights are dynamically adjusted based on the component parameters;

[0043] By collecting feedback images and conducting quality assessments, the weights are fine-tuned and the exposure time for composite scenes is recalculated in a closed-loop iterative manner.

[0044] Preferably, the quality assessment of the feedback image includes calculating the blurriness and positional offset of the feedback image;

[0045] Furthermore, the weight fine-tuning rule in the closed-loop iteration is as follows: if the ambiguity is greater than a preset ambiguity threshold, the continuous vibration compensation weight is increased; if the position offset is greater than a preset offset threshold, the impact compensation weight is increased.

[0046] Preferably, the specific implementation of calling the corresponding image processing algorithm based on the scene identifier is as follows:

[0047] If the scene is identified as the first standard scene, then the edge detection algorithm is invoked;

[0048] If the scene is identified as the second standard scene, then the line and circle detection algorithm is invoked;

[0049] If the scene is identified as a third standard scene, then the scale-invariant feature transformation algorithm is invoked.

[0050] Preferably, the method further includes a scene-related fallback mechanism, which is triggered during the size calculation process, the fallback mechanism including:

[0051] Continuously monitor the integrity of features and enable different loss rate judgment thresholds based on scene identifiers;

[0052] When the feature loss rate exceeds its corresponding judgment threshold, a rollback is triggered, and a dedicated rollback strategy bound to the corresponding scenario is executed.

[0053] After executing the rollback strategy, the validity of the restored features is verified. If the verification passes, the dimensions are recalculated based on the restored features; if the verification fails, the rollback strategy is executed repeatedly.

[0054] The beneficial effects of this invention are:

[0055] This invention decouples the combined effects of continuous vibration and transient impact on the imaging process by constructing a control framework that includes scene pre-identification, adaptive spectrum decoupling, branching compensation processing, closed-loop parameter fusion, scene association detection algorithms, and scene-based backoff mechanisms. It also matches optimal technical response strategies for different disturbance components. This fundamentally solves the problem of insufficient detection accuracy and stability under complex working conditions caused by a single processing model, enabling stable and accurate online detection of hardware component dimensions in high-speed, high-vibration industrial production environments. Attached Figure Description

[0056] Figure 1 A block diagram of the overall structure of the system of the present invention;

[0057] Figure 2 A schematic diagram of the physical layout of the system of the present invention on an industrial production line;

[0058] Figure 3 A schematic diagram of the overall process of the method of the present invention;

[0059] Figure 4 Schematic diagram of the spectral characteristics of the three standard perturbation scenarios of this invention;

[0060] Figure 5 A detailed flowchart of the parameter fusion operation performed in the third scenario according to the present invention;

[0061] Figure 6 A schematic diagram of the scenario-related fallback mechanism of this invention;

[0062] Figure 7 The internal structure block diagram of the data processing and control module of this invention. Detailed Implementation

[0063] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0064] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0065] Example 1

[0066] This embodiment discloses a vision-based inspection system for the size detection of hardware accessories, such as... Figure 2 As shown, the system of the present invention is deployed on a standard industrial production line, which includes a conveyor belt with an adjustable operating speed. The main structure of the system is securely mounted above the conveyor belt via a gantry-type support frame, ensuring the structural rigidity of the system and a low natural vibration frequency.

[0067] The overall system structure is as follows Figure 1 As shown, it mainly consists of four core parts: vibration sensing module, image acquisition module, data processing and control module, and storage module.

[0068] The vibration sensing module is used to capture mechanical vibration information of the system's environment in real time. This module is physically mounted on the camera mounting base adjacent to the image acquisition module. This mounting position most accurately reflects the vibration state transmitted to the core imaging components. The module integrates a triaxial microelectromechanical system (MEMS) accelerometer with an operating bandwidth covering low-frequency equipment vibrations from 0.5 Hz to high-frequency structural resonances up to 10 kHz, sufficient to cover most vibration sources in industrial environments. The analog voltage signal output by the sensor is connected to an analog-to-digital converter. The converted discrete digital vibration signal sequence is transmitted in real time via a serial peripheral interface (SPI) in the form of data packets to the data processing and control module for subsequent analysis.

[0069] The image acquisition module is responsible for capturing high-quality images of hardware components in dynamic environments. At its core is a CMOS image sensor with a global shutter function, effectively avoiding the rolling shutter effect caused by high-speed object movement. To eliminate perspective errors caused by slight height fluctuations of the workpiece on the conveyor belt, thus ensuring the accuracy of dimensional measurements, the module is equipped with a telecentric lens. The module also integrates a strobe LED illumination unit. This unit, arranged in a ring, consists of an array of multiple high-brightness white LEDs, capable of emitting high-intensity light pulses for an extremely short time (e.g., 10 to 100 microseconds). The flash duration of this illumination unit is synchronized with the exposure time of the CMOS image sensor via hardware circuitry. Its synchronization trigger signal is directly generated by the Field Programmable Gate Array (FPGA) 32 within the data processing and control module, ensuring a trigger delay of less than 1 microsecond, thereby achieving instantaneous frozen imaging of moving objects.

[0070] The storage module is a non-volatile memory, such as a solid-state drive, to ensure high-speed data read and write and data persistence in the event of power failure. Internally, this module stores a priori knowledge base, specifically including a historical scenario database and a closely related set of compensation rules. This historical scenario database is built upon extensive field data collection and offline analysis, and pre-defined and stored spectral characteristic models for at least three standard disturbance scenarios. (Refer to...) Figure 4The first standard scenario corresponds to continuous steady-state vibration caused by large rotating equipment (such as fans and pumps) on a production line. Its spectral characteristic model is quantified as follows: In the Fast Fourier Transform (FFT) spectrum, there is a significant dominant frequency peak in the low-frequency range of 5 Hz to 20 Hz, and the vibration displacement amplitude obtained by integrating this frequency peak is within the range of 0.1 mm to 0.5 mm. The second standard scenario corresponds to transient impacts caused by events such as conveyor belt seams passing through rollers or material falling. Its spectral characteristic model is quantified as follows: In the spectrum, there is a wideband pulse peak in the high-frequency range greater than 100 Hz. The peak acceleration corresponding to this pulse peak in the time domain is greater than 2g, and the pulse's half-width at half-maximum duration is between 10 ms and 50 ms. The third standard scenario is defined as a composite disturbance situation where the characteristics of the above two scenarios coexist. A set of parameterized compensation rules is stored corresponding to each of these three scenario models. The first compensation rule set, corresponding to the first standard scenario, is a two-dimensional lookup table. This table uses the dominant vibration frequency (1 Hz resolution) and amplitude (0.05 mm resolution) as indexes to directly map an experimentally calibrated base exposure time value. The second compensation rule set, corresponding to the second standard scenario, includes a conveyor belt seam impact feature library. This library stores standard impact waveform templates generated when the conveyor belt passes the drive roller at different speeds, used for accurate identification of impact events. It also includes an advance exposure time threshold set to 20 milliseconds. The third compensation rule set, corresponding to the third standard scenario, not only includes all the contents of the first and second compensation rule sets but also defines an additional set of initial weight values ​​for parameter fusion. For example, the initial value for continuous vibration compensation is set to 0.4, and the initial value for impact compensation is set to 0.6.

[0071] like Figure 7 As shown, the data processing and control module adopts a heterogeneous computing scheme in its hardware architecture, integrating a central processing unit (CPU) 31 responsible for high-level logic and complex algorithm operations and a field-programmable gate array (FPGA) 32 responsible for real-time timing control. The CPU and FPGA exchange data at high speed through the PCIe bus.

[0072] Example 2

[0073] This embodiment discloses a method for measuring the dimensions of hardware accessories based on vision inspection, which is implemented through the aforementioned system hardware. The detailed process is as follows: Figure 3 As shown, the specific operation steps are as follows:

[0074] First, the system performs scene pre-identification and rule loading operations. This process is executed cyclically with a fixed cycle of 100 milliseconds. The FPGA firmware logic in the data processing and control module continuously receives a 20 kHz digital vibration signal stream from the ADC of the vibration sensing module. Every time the digital signal processor (DSP) unit inside the FPGA accumulates 2000 sampling points (i.e., 100 milliseconds of data), it is treated as a data frame. Subsequently, a hardware-implemented, pipelined FFT processing core inside the FPGA performs a Fast Fourier Transform on the data frame, generating a real-time vibration spectrum within microseconds. This spectrum data is efficiently transferred to the CPU's memory via a Direct Memory Access (DMA) channel. Upon receiving new spectrum data, the CPU immediately executes a template matching algorithm, such as calculating the normalized cross-correlation coefficient between the real-time spectrum and the spectrum feature models of each standard scene in the historical scene database in the storage module. If the cross-correlation coefficient matches the first standard scene model with the highest degree of matching and exceeds a preset threshold (e.g., 0.9), the CPU identifies the current scene as the first scene. Similarly, if the match with the second standard scenario model is the highest, it is identified as the second scenario. If the match with both the first and second standard scenario models exceeds the threshold, it is identified as the third scenario. After the scenario is identified, the CPU immediately reads the compensation rule set corresponding to that scenario from the storage module and loads it into the CPU's own cache so that subsequent processing can be called with extremely low latency.

[0075] Secondly, the system performs scene-adaptive spectral decoupling. The purpose of this step is to separate the vibration components that have different impacts on imaging from the raw, mixed vibration signal. The CPU of the data processing and control module dynamically selects and executes a wavelet transform decomposition strategy based on the scene type identified in the previous step. If the current scene is the first scene (dominated by continuous vibration), the CPU calls a low-scale wavelet decomposition strategy, using the Daubechies4 (db4) mother wavelet with good time-frequency localization characteristics to decompose a recently acquired digital vibration signal sequence (e.g., 1024 points) to the 4th scale. After decomposition, only the approximate components representing the low-frequency trend of the signal are reconstructed to obtain the pure continuous vibration component. If the current scene is the second scene (dominated by transient impact), the CPU calls a high-scale wavelet decomposition strategy, also using the db4 mother wavelet, but deepening the decomposition scale to the 8th scale, and only reconstructing the detail components representing high-frequency abrupt changes in the signal to extract the transient impact component. If the current scene is the third scene, the CPU executes a full-scale decomposition strategy, decomposing to the 8th scale, and then reconstructs the approximate components in the low-frequency band and the detailed components in the high-frequency band, thereby simultaneously obtaining the continuous vibration components and the transient impact components. After the component reconstruction is completed, the CPU also performs a rapid verification of the validity of the components to eliminate noise interference: for the reconstructed continuous vibration components, it verifies whether the main frequency is below 20 Hz and whether the amplitude is less than or equal to 0.5 mm; for the reconstructed transient impact components, it verifies whether the peak acceleration is greater than or equal to 2g and whether the pulse duration is less than or equal to 50 milliseconds. After verification, to allow for sufficient margin in subsequent compensation calculations, the CPU performs scenario enhancement calculations on the component parameters: In the first scenario, the amplitude value of the continuous vibration component is multiplied by an enhancement coefficient k1 (k1=1.2), which is determined by the optimal value obtained when ambiguity is reduced by 20% in the vibration table test; In the second scenario, the peak acceleration of the transient impact component is multiplied by an enhancement coefficient k2 (k2=1.5), which is derived from the linear fitting model of the impact peak and position offset; In the third scenario, the amplitude value of the continuous vibration component is multiplied by an enhancement coefficient k3 (k3=1.1), and the peak acceleration of the transient impact component is multiplied by an enhancement coefficient k4 (k4=1.3), the values ​​of enhancement coefficients k3 and k4 are derived from the sensitivity analysis of weighted fusion under composite disturbances. The enhanced component parameters will be used in subsequent compensation calculations.

[0076] Next, the system performs scene-bound branching processing. Based on the current scene identifier, the CPU and FPGA of the data processing and control module collaboratively execute distinct processing flows to address different disturbances. In the first scene, the CPU uses the amplitude and dominant frequency of the enhanced continuous vibration component as two input parameters, queries a two-dimensional lookup table in the first compensation rule set loaded into the cache, and directly obtains a corresponding base exposure time. Simultaneously, the CPU's internal logic determines whether the enhanced amplitude exceeds a preset threshold that significantly affects image blur, such as 0.3 mm. If it exceeds this threshold, a 20% reduction operation is performed on the base exposure time obtained from the lookup table, generating a continuous vibration compensation exposure time to suppress motion blur. This exposure time value is then sent to the FPGA. In the second scene, the focus shifts to avoiding positional jumps caused by impacts. The CPU first matches the waveform of the enhanced transient impact component with the seam impact feature library in the second compensation rule set to confirm that this is a predictable impact event caused by a conveyor belt seam. Meanwhile, the FPGA uses its high-precision clock to record the exact timestamp of the impact event. Combined with real-time pulse data acquired from a rotary encoder coaxially connected to the conveyor belt drive shaft (used to calculate the instantaneous speed of the conveyor belt), the FPGA predicts the time required for the impacted metal component to reach the center of the image acquisition module's field of view using a preset linear motion model (i.e., arrival time = (camera center position - current workpiece position) / current speed). Based on this prediction, the FPGA calculates an advance exposure trigger time, set 20 milliseconds before the predicted arrival time, thus cleverly avoiding the positional jitter peak caused by the impact. In the third scenario, the system enters parallel processing mode, with the CPU and FPGA simultaneously executing the processing logic of the first and second scenarios, generating a continuous vibration compensation exposure time and an impact compensation advance trigger time, respectively.

[0077] Furthermore, the system performs scene-driven parameter fusion operations under specific conditions. This operation is triggered only when the system identifies the current scene as a third scene, and its detailed process is as follows: Figure 5As shown, the CPU of the data processing and control module first uses the initial weight values ​​defined in the loaded third compensation rule set (i.e., continuous vibration weight is 0.4, and impact weight is 0.6) to perform weighted fusion of the exposure parameters (exposure time itself and exposure timing) implied in the parallel-calculated continuous vibration compensation exposure time and impact compensation early trigger time. This fusion is not a simple numerical average, but a comprehensive decision-making process used to generate an initial composite scene exposure time. Next, the CPU dynamically adjusts the fusion weights based on the decoupled and enhanced component parameters from this loop: if the enhanced amplitude of the continuous vibration component is greater than 0.4 mm, meaning motion blur is the primary issue, the continuous vibration compensation weight is increased from 0.4 to 0.5; if the enhanced peak acceleration of the transient impact component is greater than 4g, meaning the risk of positional jump is extremely high, the impact compensation weight is increased from 0.6 to 0.7. Subsequently, the CPU recalculates the weighted fusion using the adjusted weights to obtain a more reasonable and optimized composite scene exposure time. To verify the effectiveness of these parameters, the system enters a rapid closed-loop feedback iteration process. The data processing and control module immediately controls the image acquisition module to acquire a feedback image frame according to the optimized exposure time. The image data is transmitted to the CPU, which performs a rapid quality assessment. The assessment includes two dimensions: first, quantifying the blurriness of the image by calculating the variance of the Laplacian operator of the image; the larger the variance, the clearer the image. Second, quantifying the positional offset of the image by performing phase correlation operations between the feedback image and a standard template image. If the difference between the quantified blurriness value and the ideal value is greater than a preset 10% threshold, the CPU will further slightly increase the continuous vibration compensation weight; if the positional offset is greater than 0.1 mm, the impact compensation weight will be further slightly increased. This process of weight adjustment, re-fusion calculation, acquisition of feedback images, and quality assessment is iteratively executed until all image quality indicators meet the preset requirements (e.g., blurriness index within 5% of the ideal value, positional offset less than 0.05 mm), or the maximum number of iterations is reached (e.g., 3 times). The exposure time obtained at this point is determined as the final compensated exposure time and sent to the FPGA.

[0078] Then, the system performs scene-adaptive size detection. The FPGA of the data processing and control module, based on the final exposure parameters or trigger timing received from the CPU and calculated for different scenes, sends a synchronous trigger signal accurate to the nanosecond level to the CMOS sensor and illumination unit of the image acquisition module, completing high-quality image acquisition of the hardware components on the conveyor belt. The acquired raw image data, along with its corresponding scene identifier (first, second, or third scene), is transmitted to the CPU as metadata in a data packet. The CPU parses the scene identifier in the metadata and, based on this, calls the corresponding optimized algorithm from its internal image processing algorithm library to perform feature extraction and size calculation. Specifically, if the scene identifier is the first scene, the image may have slight motion blur, and the CPU will call an edge detection algorithm based on the Canny operator. The internal parameters of this algorithm, such as the size of the Gaussian filter kernel, are set relatively large (e.g., 7×7), and the high threshold in the dual threshold is appropriately lowered to smooth noise while preserving the faint outline edges of the components that have become blurred to the maximum extent. If the scene is identified as the second scene, the main problem with the image might be macroscopic positional offset. In this case, the CPU calls a line and circle detection algorithm based on Hough transform. This algorithm is invariant to object translation and can locate key geometric references on the parts (such as the edge lines of bolts and the inner and outer circles of nuts), and use this as a basis to correct the coordinate system of the parts. If the scene is identified as the third scene, the image may have both blur and positional offset, or even slight perspective distortion. In this case, the CPU calls the Scale Invariant Feature Transform (SIFT) algorithm. This algorithm extracts locally invariant feature points in the image that are insensitive to changes in scale, rotation, and brightness. By matching these feature points with feature points in a pre-stored standard template image of the parts, a homography matrix that describes the geometric transformations between the images is calculated. Using this matrix to perform an inverse transform on the acquired image, positional offset and local distortion can be corrected simultaneously, achieving high-precision feature localization. After completing feature extraction and correction in their respective scenarios, the CPU calculates the key dimensions of the hardware accessories (such as length, outer diameter, pitch, etc.) in the corrected image coordinate system, and outputs the calculation results along with the scene identifier used in this inspection to the result database or the host computer monitoring system.

[0079] Finally, the system also includes a scene-related rollback mechanism, as shown in the attached figure. Figure 6As shown, this is designed to handle situations where image quality deteriorates significantly under extreme conditions, causing conventional algorithms to fail. During feature extraction for size detection, the CPU of the data processing and control module continuously monitors the integrity of the features and activates different judgment thresholds based on the scene identifier. In the first scene, if the total length of the detected edge features has a loss rate exceeding 15% compared to the total edge length of the standard template, backoff is triggered. In the second scene, if the Hough transform fails to detect a preset number of geometric reference features (e.g., a hexagonal nut should have six sides, but if less than five are detected), i.e., the loss rate of geometric reference features exceeds 10%, backoff is triggered. In the third scene, if the number of successful SIFT feature point matches is less than 80% of the preset total (i.e., the feature loss rate exceeds 20%), backoff is triggered. Once backoff is triggered, the CPU will execute a specific backoff strategy bound to the current scene. For the first scene, a multi-frame averaging fusion strategy is implemented. This involves retrieving the five most recent images, also identified as belonging to the first scene, stored in a memory circular buffer. These images are then precisely registered at the sub-pixel level through feature point matching, followed by a weighted average at the pixel level. This effectively suppresses random noise and "averages out" random blurring caused by vibration, thus restoring clearer edge details. For the second scene, a single-frame re-capture strategy is implemented. The CPU immediately sends a re-capture command to the FPGA. Taking advantage of the transient nature of impact events, a clear image is re-acquired within a brief stabilization period after the impact energy dissipates (e.g., 100 milliseconds after the impact). For the third scene, a three-frame feature fusion strategy is implemented. This involves retrieving the three most recent images, all identified as belonging to the third scene. Relatively clear edge features are extracted from the first frame, relatively accurate baseline features from the second frame, and the most robust keypoint features from the third frame. Finally, these three sets of features from different times and with different emphases are fused in a unified feature space to reconstruct the complete component outline in the most probabilistic way. After the rollback strategy is executed, the CPU verifies the validity of the recovered features. In the first scenario, the recovered edge sharpness index (e.g., average gradient magnitude) must be greater than 90% of the preset benchmark; in the second scenario, the positional offset of the re-captured image must be less than 0.05 mm; and in the third scenario, the overall feature extraction rate after fusion must be greater than 85%. If the verification passes, the size is recalculated based on the recovered features; if the verification fails, the corresponding rollback strategy is repeated, up to a maximum of three times. If the process fails after three attempts, the system determines that the workpiece cannot be reliably automatically inspected and sends an alarm signal requiring manual re-inspection to the external production line monitoring system or human-machine interface.

[0080] Example 3

[0081] To verify the actual effect of the technical solution of this invention, the following experimental platform was built. The conveyor belt speed was set to 1.2 m / s. The objects to be tested were a batch of M8×25 hexagonal bolts, with nominal dimensions of 25.00 mm in length and 13.00 mm in width across sides. The hardware configuration of the testing system was completely consistent with the description in the aforementioned specific embodiment. In the experiment, a typical third scenario (composite disturbance) was artificially simulated by deploying a frequency-adjustable amplitude vibration table and a pneumatic impact hammer near the system support frame: a continuous sinusoidal vibration with a frequency of 15 Hz and an amplitude of 0.4 mm was applied, while a transient impact with a peak acceleration of 5g was generated every 2 seconds.

[0082] After system startup, the data processing and control module 30 accurately identified the current environment as the third scene within 100 milliseconds through FFT spectrum analysis and loaded the third compensation rule set from the storage module. Next, the CPU performed wavelet transform, successfully decoupling the 15 Hz continuous vibration component and the 5g impact component, and enhanced their parameters. The system entered a parallel processing branch, calculating the compensation exposure time (e.g., 35 microseconds) for continuous vibration and the advance triggering timing for impact. Subsequently, the parameter fusion process began. Because the enhanced vibration amplitude (0.4mm*1.1=0.44mm) was greater than 0.4mm, and the enhanced impact peak value (5g*1.3=6.5g) was greater than 4g, the CPU increased the continuous vibration compensation weight from 0.4 to 0.5 and the impact compensation weight from 0.6 to 0.7. Based on the new weighted fusion parameters, the system performed a closed-loop feedback iteration, acquiring feedback images and performing quality assessment, ultimately determining an optimal composite scene exposure time of 38 microseconds, combined with the advance triggering timing. The FPGA controls the image acquisition module 20 to take pictures based on these final parameters. After receiving the image, the CPU, recognizing the scene as the third scene, calls the SIFT algorithm for feature matching and homography matrix correction, ultimately measuring the bolt length and opposite side width. During the continuous inspection of a large number of bolts, the system consistently operates stably in this mode, dynamically adjusting parameters to cope with each disturbance.

[0083] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A vision-based inspection system for the dimension inspection of hardware accessories, characterized in that, The system includes: The vibration sensing module is used to collect vibration signals from the system's support structure in real time. The image acquisition module is used to acquire images of hardware accessories; The storage module pre-stores a historical scenario database and an associated set of compensation rules. The historical scenario database defines spectral characteristic models for three standard disturbance scenarios, specifically: The first standard scenario corresponds to continuous steady-state vibration, and its spectral characteristic model is defined as the presence of a dominant frequency peak in the low-frequency range. The second standard scenario, corresponding to transient impact, is defined by its spectral characteristic model as the presence of a broadband pulse peak in the high-frequency range. The third standard scenario, corresponding to the composite disturbance, has its spectral characteristic model defined as the spectral characteristics of both the first and second standard scenarios. The compensation rule set stored in the storage module includes a compensation rule set that corresponds one-to-one with the three standard disturbance scenarios; A data processing and control module, which is electrically connected to the vibration sensing module, the image acquisition module, and the storage module, is configured as follows: Receive vibration signals and perform spectrum analysis to generate real-time vibration spectrum; The real-time vibration spectrum is matched with the spectral feature models of various standard disturbance scenarios in the historical scenario database to identify the type of disturbance scenario to which the current working condition belongs. Based on the identified disturbance scenario type, load the corresponding compensation rule set; Based on the loaded compensation rule set, image acquisition control parameters are generated for the current disturbance scene type; and Based on the image acquisition control parameters, the image acquisition module is controlled to perform image acquisition and process the acquired image to calculate its size.

2. The system according to claim 1, characterized in that, The vibration sensing module includes a triaxial microelectromechanical system accelerometer and an analog-to-digital converter. The accelerometer is mounted on the support structure of the image acquisition module, and the analog-to-digital converter converts the analog voltage signal into a digital vibration signal sequence.

3. The system according to claim 2, characterized in that, The data processing and control module is further configured to, after identifying the type of disturbance scenario, perform a spectrum decoupling operation based on scenario adaptation, the operation including: If the current scene is identified as the first standard scene, the low-scale wavelet decomposition strategy is invoked to decompose the digital vibration signal sequence and reconstruct the continuous vibration components. If the current scene is identified as the second standard scene, the high-scale wavelet decomposition strategy is invoked to decompose and reconstruct the transient impact component. If the current scene is identified as the third standard scene, then the full-scale decomposition strategy is executed to reconstruct the continuous vibration component and the transient impact component. Furthermore, the data processing and control module is further configured to perform scene enhancement calculations after validating the reconstructed components.

4. The system according to claim 3, characterized in that, The data processing and control module is configured to perform scene-bound branching processing operations to generate image acquisition control parameters, the operations including: If the current scene is the first standard scene, then based on the enhanced continuous vibration component, the lookup table in the first compensation rule set is queried to obtain the basic exposure time, and the basic exposure time is adjusted according to the case where the amplitude exceeds the preset threshold. If the current scene is the second standard scene, the enhanced transient impact component will be matched with the impact feature library in the second compensation rule set, the timestamp of the impact event will be recorded, and the early exposure trigger time point will be calculated. If the current scene is the third standard scene, the processing logic of the first and second standard scenes above will be executed in parallel to generate the continuous vibration compensation exposure time and the impact compensation early trigger time point, respectively.

5. A method for detecting the dimensions of hardware accessories based on vision inspection, applicable to the system described in any one of claims 1-4, characterized in that, The method includes the following steps: The vibration signal of the system support structure is collected in real time by the vibration sensing module and converted into a digital vibration signal sequence. Spectral analysis is performed on the digital vibration signal sequence, and the obtained real-time vibration spectrum is matched with the standard scene spectral feature model in the pre-stored historical scene database. This identifies the current disturbance environment as one of the first standard scene, the second standard scene, or the third standard scene, and loads the compensation rule set corresponding to the identified scene. Based on the identified scene type, an appropriate wavelet transform decomposition strategy is selected and executed to decouple the continuous vibration component and / or transient impact component from the digital vibration signal sequence, and the parameters of the decoupled components are verified and scene enhancement calculations are performed. Depending on the scene type, different processing branches are entered to calculate image acquisition control parameters; Based on the calculated image acquisition control parameters, the image acquisition module is controlled to acquire images of hardware accessories, and scene identifiers are added to the image metadata; Based on the scene identifier in the image metadata, the corresponding image processing algorithm is invoked to perform feature extraction and size calculation, and the results are output.

6. The method according to claim 5, characterized in that, In the step of calculating image acquisition control parameters, when the current disturbance environment is identified as a third standard scene, the method further includes a scene-driven parameter fusion operation, which includes: Using preset continuous vibration compensation weights and shock compensation weights, the exposure parameters implied by the parallel calculated continuous vibration compensation exposure time and the shock compensation early trigger time are weighted and fused. The weights are dynamically adjusted based on the component parameters; By collecting feedback images and conducting quality assessments, the weights are fine-tuned and the exposure time for composite scenes is recalculated in a closed-loop iterative manner.

7. The method according to claim 6, characterized in that, The quality assessment of the feedback image includes calculating the blur and position offset of the feedback image; Furthermore, the weight fine-tuning rule in the closed-loop iteration is as follows: if the ambiguity is greater than a preset ambiguity threshold, the continuous vibration compensation weight is increased; if the position offset is greater than a preset offset threshold, the impact compensation weight is increased.

8. The method according to claim 7, characterized in that, The specific implementation of calling the corresponding image processing algorithm based on the scene identifier is as follows: If the scene is identified as the first standard scene, then the edge detection algorithm is invoked; If the scene is identified as the second standard scene, then the line and circle detection algorithm is invoked; If the scene is identified as a third standard scene, then the scale-invariant feature transformation algorithm is invoked.

9. The method according to claim 8, characterized in that, The method also includes a scene-related fallback mechanism, which is triggered during size calculation. The fallback mechanism includes: Continuously monitor the integrity of features and enable different loss rate judgment thresholds based on scene identifiers; When the feature loss rate exceeds its corresponding judgment threshold, a rollback is triggered, and a dedicated rollback strategy bound to the corresponding scenario is executed. After executing the rollback strategy, the validity of the restored features is verified. If the verification passes, the dimensions are recalculated based on the restored features; if the verification fails, the rollback strategy is executed repeatedly.

Citation Information

Patent Citations

  • A metal stamping detection method and system based on machine vision

    CN118840312B

  • Alignment compensation method and system for edge microcrack detection, medium and product

    CN119043171A

  • Position sensor for image blur correcting optical system

    JP1998062831A