A visual inspection system for surface defects of mechanical parts
By combining a modular system with optical imaging, geometric shape perception, and dynamic polarization control, the problems of blind spots and ambient light sensitivity in traditional detection methods on complex curved and highly reflective parts are solved, achieving high-precision defect detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-08
- Publication Date
- 2026-03-24
AI Technical Summary
Traditional visual inspection methods cannot achieve deep coupling between real-time geometric shape perception and polarization optics control when dealing with complex curved and highly reflective parts, resulting in blind spots, sensitivity to ambient light, motion blur, and insufficient detection capabilities for minute defects.
By employing an optical imaging module, a geometric shape perception module, a dynamic polarization control module, and an image fusion and enhancement module, the polarization state is dynamically adjusted through real-time calculation of three-dimensional geometric information. Combined with multi-frame image fusion and feature enhancement processing, high-precision defect detection of complex curved surface parts is achieved.
It achieves fully automated, high-precision, blind-spot-free inspection of complex curved surface parts, reduces the system's sensitivity to changes in ambient light, ensures imaging stability and consistency, and significantly improves the ability to detect minute defects.
Smart Images

Figure CN121482043B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of machine vision detection, and particularly relates to a mechanical part surface defect vision detection system. BACKGROUND
[0002] In industrial automation and intelligent manufacturing, the surface quality of mechanical parts is directly related to the reliability and life of products, especially those metal parts with complex curved surfaces and strong surface reflection, such as turbine blades, bearing rings, etc. The defect detection of these parts has always been a technical difficulty. When facing such parts, the traditional visual method is easily disturbed by factors such as mirror reflection and curved surface changes, resulting in unstable detection and high missed detection rate.
[0003] At present, the detection method based on polarization imaging is used to a certain extent to suppress reflection and enhance defect contrast. The commonly used technology is to use fixed orthogonal polarization structure or static polarization parameter optimization strategy based on preset threshold. However, this method has some obvious defects in actual production line: first, the fixed polarization setting is difficult to follow the curvature change of the curved surface, and it is easy to form a blind area at the corner or curved part of the part. The preset parameter is sensitive to changes in light environment, and the imaging effect may decrease significantly when the workshop light changes slightly, and the false positive rate increases. Second, the existing system mostly separates the "shape" and "polarization" processing, that is, first knowing what the part looks like, and then imaging according to the fixed mode, which leads to the fact that in dynamic detection, especially when the mechanical arm moves with the camera, the polarization state cannot be quickly adjusted according to the real-time curved surface shape, which is easy to produce motion blur and slow down the detection rhythm. In addition, due to the limitations of camera resolution and optical diffraction, the detection ability of the existing method for fine scratches and pits is still limited.
[0004] In view of the above problems, the application provides a solution. SUMMARY
[0005] The application aims to provide a mechanical part surface defect vision detection system to solve the problems of detection blind area, environment light sensitivity, motion blur and insufficient detection ability of micro defects caused by the inability to realize the deep coupling of real-time geometric shape perception and polarization optical control when the fixed polarization structure or static polarization optimization strategy in the prior art faces complex curved surface high reflection parts.
[0006] The technical scheme of the application is a mechanical part surface defect vision detection system, comprising:
[0007] An optical imaging module for acquiring original image data of the surface of the mechanical part to be detected;
[0008] A geometric shape perception module coaxially integrated with the optical imaging module for real-time solving of three-dimensional geometric information of the surface of the part to be detected;
[0009] The dynamic polarization control module, connected to the geometry perception module and the optical imaging module, is used to dynamically adjust the polarization state of the illumination light incident on the surface of the part and the polarization state of the reflected light received by the camera based on the three-dimensional geometric information output in real time by the geometry perception module.
[0010] The image fusion and enhancement module, connected to the optical imaging module and the dynamic polarization control module, is used to fuse and enhance the features of at least three multi-frame polarization images acquired under different polarization states controlled by the dynamic polarization control module to obtain an enhanced image.
[0011] The defect identification and decision-making module, connected to the image fusion and enhancement module, is used to automatically detect and classify defects in enhanced images and output the detection results.
[0012] Preferably, the optical imaging module includes a high-resolution global shutter industrial camera, a fixed-focus telecentric lens, and a controllable illumination unit;
[0013] The controllable lighting unit is a multi-layer ring array of multi-angle LED light sources, containing at least two coaxial and independently controlled lighting sub-rings. The angle between the luminous surface normal of each lighting sub-ring and the optical axis of the system is different, and the brightness can be adjusted independently and continuously.
[0014] Preferably, the geometric shape perception module merges the structured light projection path with the camera imaging path through a beam splitter prism;
[0015] The geometry perception module includes a structured light projection unit for projecting a set of spatially and temporally encoded light stripe patterns onto the surface of the part.
[0016] A high-resolution global shutter industrial camera synchronously acquires images of deformed stripes modulated by the three-dimensional shape of the part surface;
[0017] The geometry perception module also includes a depth calculation unit, which processes the deformed stripe image, reconstructs the three-dimensional point cloud of the part surface in real time, and estimates the surface normal direction.
[0018] Preferably, the dynamic polarization control module includes an incident polarization control unit and a receiving polarization control unit;
[0019] The incident polarization control unit is located at the optical path exit of the controllable illumination unit, and includes a linear polarizer and a first drive mechanism for driving its rotation.
[0020] The receiving polarization control unit is located at the front end of the industrial camera lens and includes an analyzer and a second drive mechanism that drives its rotation.
[0021] Preferably, the dynamic polarization control module has an embedded polarization optimization model. The polarization optimization model takes the surface normal vector, incident light direction vector, camera observation direction vector and optical constants of the part material from the geometric shape perception module as input, and takes maximizing the imaging signal contrast between the defect area and the normal background area as the optimization objective. It calculates the optimal incident polarization angle and the optimal received polarization angle for each pixel in parallel.
[0022] Preferably, the image fusion and enhancement module performs the following operations:
[0023] Subpixel-level registration of multi-frame polarization images;
[0024] Calculate the difference images between each pair of registered images;
[0025] Based on the theoretical contrast gain value calculated and stored for each image pixel by the dynamic polarization control module, adaptive fusion weights are assigned to each original image and difference image and weighted fusion is performed to obtain a preliminary fused image.
[0026] Local contrast enhancement processing is performed on the initially fused image.
[0027] Preferably, the image fusion and enhancement module also performs diffraction enhancement operations for subpixel-level minute defects:
[0028] Based on the depth information of each pixel provided by the geometric shape perception module and the optical parameters of the imaging system, digital refocusing is performed using scalar diffraction theory to remove the blurred details caused by complex optical diffraction and defocusing effects.
[0029] Preferably, the defect recognition and decision-making module includes a fully convolutional network model with an encoder-decoder structure, used to perform pixel-level defect probability prediction on the enhanced image and generate a defect segmentation probability map;
[0030] Based on the defect segmentation probability map, binarization, connected component analysis, feature extraction and classification are performed to generate a structured inspection report.
[0031] Preferably, the geometric shape perception module, the dynamic polarization control module, and the optical imaging module are connected in series to form a millisecond-level polarization fast response closed loop, whose working cycle is synchronized with the acquisition frame rate of the high-resolution global shutter industrial camera.
[0032] The output of the defect identification and decision-making module is connected to a second-level system parameter adaptive loop, which is used to slowly optimize the parameters in the polarization optimization model or the image fusion and enhancement module based on historical detection performance statistics.
[0033] Preferably, the dynamic polarization control module is also used for:
[0034] Cluster analysis is performed on the surface normal vectors of each pixel in the current image output by the geometric shape perception module. Pixels with similar normal directions are divided into the same control region. The optimal incident polarization angle and the optimal received polarization angle are calculated for each control region. Then, the optimal polarization angle combination command for each region is generated and sent to the incident polarization control unit and the received polarization control unit for execution.
[0035] In summary, this application includes at least one of the following beneficial technical effects:
[0036] 1. This invention acquires the three-dimensional information of the part surface in real time through a geometric shape perception module, and dynamically adjusts the illumination polarization angle and camera polarization detection angle accordingly to achieve depth adaptive matching between optical imaging parameters and instantaneous surface geometry. This solves the detection blind zone problem caused by drastic changes in reflection angle on complex curved surfaces in traditional fixed polarization strategies, while reducing the system's sensitivity to changes in ambient light in the workshop and ensuring the stability and consistency of imaging during dynamic scanning.
[0037] 2. This invention uses multiple frames of images under different optimized polarization states for fusion and weights the contrast gain based on the polarization physics model prediction, which can effectively highlight the polarization modulation characteristics of defects. At the same time, combined with digital diffraction enhancement processing based on accurate three-dimensional depth information, it can effectively restore sub-pixel-level defect details blurred by optical diffraction and defocus. Attached Figure Description
[0038] Fig. 1 This is a flowchart illustrating the visual inspection system for surface defects in mechanical parts according to the present invention.
[0039] Fig. 2 This is a flowchart illustrating the process of the dynamic polarization control module of the present invention calculating the optimal polarization state in real time based on geometric topography.
[0040] Fig. 3 This is a schematic diagram of the multi-frame polarization image fusion and diffraction enhancement process of the image fusion and enhancement module of the present invention. Detailed Implementation
[0041] To further illustrate the technical means and effects adopted by the present invention in order to achieve the intended purpose, the following detailed description is provided in conjunction with the accompanying drawings and preferred embodiments, based on specific implementation methods of the present invention.
[0042] This embodiment provides a visual inspection system for surface defects of mechanical parts, which aims to perform fully automatic, high-precision, and blind-spot-free visual inspection of surface defects of mechanical parts with complex curved surfaces and highly reflective surfaces. By physically deploying it on a six-degree-of-freedom industrial robot end effector, it forms a freely movable online inspection workstation. The controller of the industrial robot establishes a real-time communication link with the core computing unit of the inspection system through gigabit Ethernet, ensuring strict synchronization and data interaction between motion control, optical perception, and image processing.
[0043] This embodiment consists of five functional modules, forming a complete closed loop from optical signal acquisition to defect decision output. The modules are optical imaging module, geometric shape perception module, dynamic polarization control module, image fusion and enhancement module, and defect recognition and decision module.
[0044] All of the above modules are integrated into an industrial-grade chassis with precise thermal design and electromagnetic shielding. The chassis is rigidly connected to the end flange of the industrial robot through high-rigidity, low-vibration connectors to ensure the stability of the imaging optical path during high-speed movement.
[0045] To further understand the relationships between the modules in this solution, the following will be combined with... Figs. 1 to 3 Each module in this embodiment will be described in detail.
[0046] The optical imaging module is used to acquire raw image data of the surface of the mechanical part to be inspected, specifically including:
[0047] S101: Employs a high-resolution global shutter industrial camera as the core imaging device.
[0048] To ensure distortion-free images during high-speed robotic arm movements, the camera must have a global shutter function to eliminate motion distortion introduced by the rolling shutter. Simultaneously, to effectively distinguish subpixel-level minute defects (such as micro-scratches), the camera sensor resolution should not be too low, typically no less than 8 megapixels, and it needs a high frame rate (generally no less than 30 frames per second) to meet the real-time requirements of dynamic acquisition.
[0049] In this embodiment, a 1-inch back-illuminated CMOS sensor camera is preferred, with an effective pixel count of 12 million (4000×3000 pixel array) and a maximum frame rate of 60 frames per second.
[0050] S102: Industrial cameras are forced to enable global shutter mode to completely eliminate rolling shutter distortion that may occur when the robot carries the camera at high speed. This can usually be set through the camera's configuration software or application programming interface, or by directly selecting camera hardware that only supports global shutter mode.
[0051] S103: Equipped with a fixed-focus telecentric lens for the camera, its characteristic is that the principal ray is parallel to the optical axis, so that the magnification of the image does not change with the slight change in the distance of the object, which is crucial for the subsequent three-dimensional geometric measurement based on the position of image pixels.
[0052] The selected telecentric lens should have low distortion (e.g., distortion coefficient less than 0.1%) and a suitable working distance (e.g., greater than 200mm) to accommodate common part sizes and installation spaces.
[0053] In this embodiment, a fixed-focus telecentric lens with a focal length of 35mm is selected.
[0054] S104: To achieve multi-angle adaptive lighting for complex curved surface parts, a multi-layer ring array of multi-angle LED light sources is configured as a controllable lighting unit. The mechanical structure of this LED light source consists of four coaxial and independently controlled lighting sub-rings, numbered from the inside to the outside as the 1st to the 4th lighting sub-rings.
[0055] S105: Each lighting sub-ring integrates multiple high-brightness white LED chips, which are evenly arranged along the circumference of the ring to ensure uniform light intensity distribution within the lighting area of the ring.
[0056] In this embodiment, each illumination sub-ring integrates 32 white LEDs with a color temperature of 6500K (close to sunlight), which can provide sufficient illumination while maintaining good color rendering, which is beneficial for subsequent image processing.
[0057] S106: The angle between the luminous surface normal of each illumination sub-ring and the system optical axis (camera optical axis) is fixed and increases sequentially from the inner ring to the outer ring, forming an angle gradient from small to large. The purpose is that when inspecting parts with complex curved surfaces, no matter how the local surface normal direction changes, there will always be one or more illumination angles that can return to the camera in a direction close to the specular reflection angle, thereby providing strong and controllable signal light for polarization control.
[0058] In this embodiment, the included angles of the four illumination sub-rings from the inside to the outside are set to 15 degrees, 30 degrees, 45 degrees and 60 degrees respectively, so as to cover the common range of curved surface variations. For example, for steep curved surface areas with a large angle between the normal and the optical axis, large-angle (such as 60 degrees) illumination is more likely to form effective specular reflection.
[0059] S107: The brightness of each lighting sub-ring can be adjusted independently and continuously. This is achieved by controlling the driving current of the sub-ring within the range of 0-500mA via a PWM signal. This allows the system to optimize the lighting intensity in real time according to the actual needs of different lighting angles and the local reflection characteristics of the part surface.
[0060] In this embodiment, to ensure that the control can keep up with the rhythm of the system's dynamic adjustment, the response time of the entire brightness adjustment loop is designed to be within 1 millisecond.
[0061] The geometric shape perception module is coaxially integrated with the optical imaging module to calculate the three-dimensional geometric information of the surface of the part to be inspected in real time.
[0062] In this embodiment, a beam splitter is used to merge the structured light projection path and the camera imaging path, ensuring that the center of the projector (projection pixel) and the optical center of the industrial camera are spatially highly aligned and share the same principal optical axis. This allows for real-time calculation of the three-dimensional geometric information of the surface of the part to be inspected. The data update frequency is strictly synchronized with the camera's frame rate, both being 60 times per second. The specific implementation process is broken down as follows:
[0063] S201: Structured light pattern projection.
[0064] The structured light projection unit of the geometry perception module is a blue light projector based on a digital micromirror device with an output wavelength of 450 nanometers. The output light intensity of the projector is calibrated to ensure that the light spot formed on the surface of the part to be inspected is uniform.
[0065] The projector projects a set of light stripe patterns that are encoded in both space and time onto the surface of the part to be inspected. In this embodiment, the encoding strategy used is a combination of Gray code and four-step phase shifting method. This can both use Gray code to determine the absolute phase order over a large range and use the four-step phase shifting method to obtain high-precision phase principal values.
[0066] The specific projection process is executed sequentially within a complete measurement cycle:
[0067] S201a: The projector sequentially projects three black-and-white binarized Gray code patterns, which are used in subsequent processing to determine which complete stripe period each pixel belongs to.
[0068] S201b: The projector quickly switches and projects four stripe patterns with sinusoidal brightness and phase shifts of 90 degrees between each other (i.e., phase shift steps of 0 degrees, 90 degrees, 180 degrees and 270 degrees respectively), which are used to calculate the fine phase value of each pixel in one cycle.
[0069] All patterns are pre-stored in the projector's memory. The projection process is controlled by a hardware trigger signal that is strictly synchronized with the exposure signal of the industrial camera, ensuring that the projection time of each structured light pattern corresponds precisely to the image acquisition time of the industrial camera.
[0070] S202: Image acquisition and transmission of deformed stripes.
[0071] While the industrial camera projects the pattern onto the projector, it simultaneously acquires images of deformed stripes modulated by the three-dimensional shape of the surface of the part to be inspected.
[0072] To ensure stable transmission of high-resolution images (such as 12 megapixels) at a rate of 60 frames per second, the acquired image data is directly transmitted to the module's dedicated depth computing unit via high-speed industrial camera interfaces such as Camera Link Full or CoaXPress.
[0073] S203: Phase Calculation and Decoding.
[0074] The depth computing unit is a dedicated processing board that integrates an FPGA (Field-Programmable Gate Array) and a GPU (Graphics Processing Unit). The two are interconnected via a high-speed bus to form a collaborative computing architecture. After receiving continuously acquired stripe images, this unit processes them according to the following process:
[0075] S203a: Package phase calculation.
[0076] The four phase-shifted images acquired simultaneously are preprocessed. Then, for each pixel in the image, the gray value of the pixel in the four images is used to calculate a principal phase value in the interval (-π, π] according to the mathematical relationship of the four-step phase-shifting method. This principal phase value reflects the precise position of the pixel in one fringe period, but there is periodic ambiguity.
[0077] S203b: Absolute phase decoding.
[0078] The periodic ambiguity problem is solved by using three previously acquired Gray code patterns. The Gray code images are binarized, and an integer sequence number (i.e., stripe order) is decoded at each pixel position, representing which complete stripe period the point belongs to.
[0079] Finally, this stripe order is combined with the phase principal value calculated in the first step to obtain an absolute phase value that is continuous and monotonically increasing from the top left corner to the bottom right corner of the image. This absolute phase value determines the position of each pixel in the overall coded pattern.
[0080] S204: 3D point cloud reconstruction and normal estimation.
[0081] After obtaining the absolute phase value, the depth calculation unit, based on the principle of triangulation, uses the parameters obtained by the system's pre-calibrated high precision to reconstruct the three-dimensional point cloud of the part surface in real time.
[0082] Specifically, during the calibration phase, the system has accurately established the mapping relationship between absolute phase values and three-dimensional spatial coordinates. For each pixel in the image, based on its absolute phase value, the depth value corresponding to that point can be directly calculated by looking up a table or using the conversion formula obtained from the calibration. Then, combined with the pixel coordinates and camera internal parameters, the complete coordinates of that point in three-dimensional space can be obtained. The complete coordinates of each pixel in three-dimensional space constitute a three-dimensional point cloud.
[0083] Subsequently, the depth computing unit processes the reconstructed dense 3D point cloud in real time, estimating the surface normal direction of the part corresponding to each pixel:
[0084] S204a: Local plane fitting.
[0085] For each 3D point, select the 3D points corresponding to its 8 neighboring pixels in the image to form a local point set. Assuming that the surface of this small region can be approximated by a plane, solve for an optimal plane equation using the least squares method. The normal vector of this plane represents the surface orientation of the local region.
[0086] S204b: Calculation and output of unit normals.
[0087] The optimal plane normal vector obtained above is normalized to obtain a unit normal vector of length 1, which is used as the surface normal output of the pixel. This process is executed in parallel on all pixels on the GPU to achieve millisecond-level speed.
[0088] The depth calculation unit has a measurement accuracy better than 0.05 mm in the depth direction and better than 0.02 mm in the planar direction; the end-to-end processing latency of the entire process from image acquisition to output of 3D point cloud and normal vector is controlled within 3 milliseconds.
[0089] The dynamic polarization control module is used to dynamically adjust the polarization state of the illumination light incident on the surface of the part to be inspected and the polarization state of the reflected light received by the camera, based on the three-dimensional geometric information output in real time by the geometric shape perception module.
[0090] The dynamic polarization control module constructs a fast optical control closed loop based on surface geometry feedback, as detailed below:
[0091] S301: Module Functions and Hardware Configuration.
[0092] The dynamic polarization control module includes an incident polarization control unit and a receiving polarization control unit, which work together to achieve joint control of the polarization state.
[0093] S301a: An incident polarization control unit is set up. This unit is physically located at the optical path exit of the controllable illumination unit (close to the front of the outermost illumination sub-ring). Its core is a linear polarizer with a high extinction ratio (extinction ratio greater than 10000:1 within the system's operating wavelength range). This linear polarizer is mounted on the rotating shaft of a high-precision closed-loop stepper motor by a mechanical clamp. The stepper motor adopts closed-loop control, and its rotor shaft end is integrated with a 24-bit absolute encoder for real-time feedback of the shaft's angular position. This design enables the electrical control resolution of the motor's rotation angle to reach 0.1 degrees, and the repeatability is better than 0.05 degrees.
[0094] The motor drive controller receives digital angle commands from the main processor of the dynamic polarization control module via fieldbus such as RS-485 or CAN. These commands typically contain the target angle value (a floating-point number in degrees). The drive controller internally calculates and generates corresponding pulse sequences and current control signals based on the absolute position fed back by the current encoder and the target position, thereby driving the motor to move along the optimal acceleration and deceleration curve.
[0095] Under the typical load (polarizer and fixture) of this embodiment, the entire process from the controller parsing the instruction to the motor rotating and stabilizing at the target angle position can be completed within 5 milliseconds.
[0096] S301b: Sets up a receiving polarization control unit, which is located at the standard filter thread at the front of the industrial camera lens for easy installation and calibration. Its core is an analyzer with a high extinction ratio, which is rigidly connected to a piezoelectric ceramic-based resonant rotary actuator.
[0097] The resonant rotary actuator utilizes the inverse piezoelectric effect of piezoelectric ceramics. A sinusoidal AC voltage with a 90-degree phase difference is applied to the piezoelectric ceramic sheets attached in two orthogonal directions, which excites the actuator to generate microscopic elliptical motion. Through friction coupling or a flexible hinge mechanism, this microscopic vibration is converted into the macroscopic rotational motion of the analyzer around the optical axis.
[0098] By precisely controlling the frequency of the drive voltage (typically in the range of hundreds of hertz) and the phase difference between the two voltages, the direction of rotation and the step angle can be precisely controlled. The driver is usually equipped with a position sensor (such as a miniature capacitive or strain sensor) to form a position closed loop to ensure the accuracy of angle control. Its control interface receives digital angle commands from the main processor, and the internal circuitry converts them into the required voltage frequency and phase control words.
[0099] The drive can rotate at speeds up to 1000 degrees per second, achieving a step resolution of 0.1 degrees. It also has an extremely short electromechanical response delay of less than 1 millisecond from the time a command is issued to the start of stable motion, thus enabling it to closely follow the camera's acquisition rhythm of up to 60 frames per second or more.
[0100] S302: Core polarization optimization model.
[0101] The dynamic polarization control module embeds a dedicated graphics processing unit (GPU) that runs a preset polarization optimization model, which is the core algorithm for achieving adaptive control. Specifically:
[0102] S302a: For each pixel (i,j) within the current imaging field of view, the model input parameters include:
[0103] Surface normal vector N from the geometry perception module ij ;
[0104] The incident light direction vector L, derived from system calibration data, for the currently active illumination sub-ring;
[0105] The camera observation direction vector V (determined after system calibration) is fixed for the image center point after system calibration. For non-center pixels, it can be calculated based on the camera lens model and pixel coordinates.
[0106] The basic optical constants of the material of the part to be tested, namely the real part n and the imaginary part k of the complex refractive index (entered for a specific material during system initialization).
[0107] S302b: The GPU performs polarization optimization calculations in parallel. The polarization optimization model is based on the polarization optics theory of the Mueller matrix and Stokes vector. It models the local part surface corresponding to each pixel in the imaging field of view as a rough surface composed of a large number of randomly oriented micro-facets. In this model:
[0108] Normal background regions are considered to conform to the statistically uniform micro-area distribution characteristics;
[0109] Defective regions (such as scratches and pits) are modeled as micro-facet distributions with different statistical properties. For example, defects may cause the orientation distribution of micro-facets to be more disordered or the surface slope distribution to change. This difference will be reflected in the specific parameters of the micro-facet bidirectional reflectance distribution function (pBRDF) model.
[0110] The core of the model is to construct and maximize an objective function that quantitatively describes the contrast between the defect region and the surrounding normal background region in the imaging signal under given incident and receiving polarization states. Specifically, this objective function is defined as:
[0111] The backscattered light mainly induced by the defect region (its polarization state is represented by the Mueller matrix M) defect (Description) and the specular reflection light mainly generated in the normal background region (whose polarization state is represented by the Mueller matrix M) background(Description) The difference in signal intensity under specific observation conditions, which is calculated by multiplying the Stokes vector of the incident light by M, the surface reflection (each multiplied by M). defect Or M background The final light intensity is calculated after passing through the analyzer.
[0112] The objective function C(α,β) is mathematically expressed based on a polarization optical transmission model. Its core is calculating the difference in imaging light intensity between the defect region and the background region under specific incident polarization angles α and β. Specifically, the function is constructed as follows:
[0113] 1. Define the Stokes vector of the incident light as S. in (α), which is determined by the state α of the incident polarization control unit.
[0114] 2. The reflection characteristics of the defective region and the normal background region are respectively determined by their Mueller matrix M. defect and M background The two matrices are described as functions of surface microgeometry, material optical constants (n, k), and incident-observation geometry (determined by normal N, optical path L, V).
[0115] 3. The Mueller matrix of the receiver analyzer is A(β), which is determined by the state β of the receiver polarization control unit.
[0116] 4. The light intensity I that finally reaches the defect area of the camera defect and background area light intensity I background They can be represented as:
[0117] ;
[0118] ;
[0119] in, This indicates taking the first element of the Stokes vector (i.e., the total light intensity), where K is a constant related to the system gain.
[0120] 5. The objective function C(α,β) is defined as the absolute value or square of the difference in light intensity between the two values mentioned above, in order to maximize the contrast:
[0121] ;
[0122] or:
[0123] ;
[0124] The optimization process is to find the optimal... The combination maximizes C(α,β). defect With M backgroundThe specific parameter differences are obtained by learning from a large number of known defect samples and normal surface polarization imaging data, or calculated based on a preset physical defect model (such as a micro-surface element model with a specific slope distribution).
[0125] The optimization calculations described above are performed in a highly parallel manner on the GPU. For a frame of 4000×3000 pixels (12 million pixels in total), the GPU starts an equal number of computation threads. Each thread is independently responsible for one pixel, and calculates the objective function value using the parameters in step S302a above and the trial value of (α,β) in the current iteration as input.
[0126] The optimization process employs gradient descent, using the partial derivatives of the objective function with respect to α and β to guide the search direction.
[0127] In practice, by setting a fixed step size (learning rate) and a criterion of stopping iteration when the change in the objective function value is less than a preset threshold, this algorithm can usually converge to a local optimum within 10 iterations, that is, calculate an optimal set of incident polarization angles α for that pixel. opt and the optimal receiving polarization angle β opt .
[0128] The entire parallel optimization calculation process for all pixels in a single frame is strictly controlled to be completed within a 2-millisecond time window.
[0129] S303: Control command generation and issuance.
[0130] After the GPU completes the calculation of the optimal polarization angle for all pixels, the dynamic polarization control module generates executable instructions and coordinates hardware actions according to the following steps:
[0131] S303a: Region clustering and instruction generation based on normal direction.
[0132] The system will not directly process 12 million individual control points.
[0133] First, the system calculates the surface normal vectors (N) of all pixels in the current image. ij Cluster analysis is performed, specifically using the K-means clustering algorithm based on Euclidean distance. The three-dimensional unit normal vector is used as a feature, and normals that are close in position in the feature space (unit sphere) are grouped into the same class.
[0134] The number of clusters is dynamically adjusted according to the complexity of the surface, but is usually limited to between 50 and 200 to ensure manageability. For example, a relatively flat surface may generate only a dozen regions, while a complex freeform surface may generate hundreds of regions.
[0135] After clustering, the system assigns a value to each region R.k To calculate an optimal combination of polarization angles representing the region, a common approach is to take the α values calculated from all pixels within the region. opt and β opt The median or mean of the value is used as the uniform command angle (α) for the region. k ,β k ).
[0136] Finally, the system generates a structured list of instructions, each item of which contains at least: region identifier k, target incident polarization angle α. k Target receiving polarization angle β k .
[0137] S303b: Instruction encapsulation and high-speed delivery.
[0138] The above instruction list is encapsulated into a data packet and sent through a low-latency communication bus (such as PCIe or Gigabit Ethernet). The sending time must be within the current frame image processing cycle, and sufficient time must be allowed for physical execution.
[0139] Specifically, the instructions need to be sent to the local controllers of the two control units at a set time point (e.g., 2-3 milliseconds in advance) before the exposure trigger signal of the next frame of the camera arrives, to ensure that the actuators have enough time to complete the movement.
[0140] S303c: Institutional Coordination Positioning and Exposure Triggering.
[0141] The local controllers of the two control units immediately begin execution upon receiving the instruction:
[0142] The incident polarization control unit, whose stepper motor controller is based on the received α k The command (corresponding to the area it is responsible for) drives the motor to rotate, and as mentioned earlier, it can be positioned to the target angle in about 5 milliseconds under typical load.
[0143] The receiving polarization control unit, whose piezoelectric ceramic actuator operates according to the received β k The command drives the analyzer to rotate, and its response is extremely fast, reaching the target angle within 1 millisecond.
[0144] In summary, after issuing a command, the main controller waits for the two control units to return a "positioning complete" signal. Only after confirming that both are stable at the target angle will the main controller send a command to the industrial camera to trigger the next frame exposure. This ensures that each frame of image acquisition is performed under a polarization state specifically optimized for the current surface geometry, thus creating optimal conditions for enhancing defect contrast from the very beginning of the imaging process.
[0145] Through the complete process described above, from sensing the geometry to calculating the optimal polarization angle, and then to executing and confirming the partition control, the dynamic polarization control module achieves rapid adaptation based on the real-time surface morphology. Regardless of whether the surface of the part is flat, curved, or complex, the polarization state of the incident light and the received light can maintain the best real-time match with the optical reflection characteristics of the local surface, thereby enhancing the contrast between the defect area and the normal background from the imaging source.
[0146] The image fusion and enhancement module is used to fuse and enhance the features of multiple frames of polarization images acquired after being controlled by the dynamic polarization adjustment module, in order to improve the signal-to-noise ratio, enhance defect contrast, and reconstruct details of minute defects. Specifically:
[0147] S401: Multi-frame image acquisition and input.
[0148] The image fusion and enhancement module begins by acquiring a specific sequence of images. The system controls the industrial camera to continuously acquire at least 3 frames of images within a pre-set, extremely short time window (e.g., less than 50 milliseconds) after receiving a unified acquisition trigger signal.
[0149] The setting of this time window is mainly based on two considerations:
[0150] First, it is much smaller than the attitude stabilization period that naturally occurs due to path planning when an industrial robot carries a camera to scan parts, thus ensuring that the relative motion between the camera and the parts can be almost ignored within this window.
[0151] Secondly, it matches the instruction refresh cycle of the dynamic polarization control module, ensuring that multiple frames of images acquired within this window can correspond to a set (at least 3) of different optimal polarization state combinations that are calculated in real time and quickly switched by the dynamic polarization control module.
[0152] Based on this, the continuously acquired images are denoted as I1, I2, and I3.
[0153] In practice, by taking the "short pause" or "slow pass" segment of the robot on the scanning path as the acquisition segment and issuing continuous trigger signals during this stage, the imaging condition of "relative displacement less than 1 pixel" can be easily met. Therefore, this set of images I1, I2, I3 not only has a high time correlation, but also carries optical information for the same surface position but under different optimized polarization states.
[0154] S402: Subpixel-level image registration.
[0155] After acquiring multiple frames of images, high-precision spatial alignment, or subpixel-level image registration, is required. Although the acquisition time window is extremely short and the relative displacement is minimal, any minute differences in translation, rotation, or deformation between images must be eliminated to ensure the accuracy of subsequent pixel-level differencing and fusion operations. The registration process involves two steps:
[0156] The first step is to roughly estimate the overall translation amount.
[0157] The goal is to quickly estimate the most significant displacement deviation between images. Using one frame (e.g., I1) as a reference, the overall translation is calculated for another frame (e.g., I2) using the phase correlation method. The specific steps are as follows:
[0158] The two images are converted to the frequency domain, the cross power spectrum is calculated, and then the cross-correlation plane is obtained through inverse Fourier transform. On this plane, the position of the peak corresponds to the integer pixel translation of image I2 relative to I1 in the horizontal and vertical directions.
[0159] To improve accuracy, surface sampling is performed on the region near the peak, thereby improving the accuracy of translation estimation to the sub-pixel level (e.g., 0.1 pixels), which can quickly correct major offsets caused by minor robot jitters or differences in trigger timing.
[0160] The second step is to refine the local nonlinear deformation.
[0161] After the first step of overall translation correction, there may still be local nonlinear deformations between images caused by small changes in viewpoint or curved surface perspective effects. In order to correct these finer deviations, hundreds of feature points are uniformly selected on the image (for example, extracted using SIFT or ORB algorithms, or points are directly selected on a regular grid). Then, between each pair of adjacent images, local deformation optimization based on gradient descent is performed on these feature points.
[0162] The goal of optimization is to minimize the positional error between all corresponding feature point pairs, assuming that the geometric transformation between two adjacent images can be approximated by a local affine transformation model (for small disparity regions).
[0163] For each local region (e.g., a small window centered on a feature point), the gradient descent algorithm iteratively adjusts the six parameters of the affine transformation to minimize the pixel intensity difference (or normalized cross-correlation value) between the two images within the window. This process is performed in parallel across all selected feature point regions.
[0164] Ultimately, the system will obtain a dense displacement field that describes the sub-pixel level precise position that each pixel in the second image needs to move relative to the first reference image. Using this displacement field, images I2, I3, etc., can be accurately resampled onto a coordinate grid that is perfectly aligned with image I1 through methods such as bilinear interpolation.
[0165] By employing the two-step strategy of overall coarse matching and local fine-tuning, it can be ensured that in all multi-frame images used for fusion, the light rays from the same physical point on the surface of the part to be detected ultimately fall on the same sub-pixel position in the fused image (the overall registration accuracy is better than 0.1 pixels).
[0166] S403: Image fusion based on polarization difference and adaptive weighting.
[0167] After registration, an image fusion algorithm that deeply integrates polarization information is executed. For a pixel in the i-th row and j-th column of the registered image, the algorithm proceeds as follows:
[0168] S403a: Calculate the absolute difference between any two images and generate three difference images:
[0169] ;
[0170] ;
[0171] .
[0172] Differential operations can effectively suppress background noise and uniform reflected light that do not change with polarization state, while highlighting polarization-sensitive features such as defect edges.
[0173] S403b: Calculate adaptive fusion weights.
[0174] The system will read the optimal polarization angle used by the dynamic polarization control module for the current pixel (i,j) when acquiring image I_k. Below, the "theoretical contrast gain value" is synchronously calculated and stored by the polarization optimization model. . It is a numerical value representing the theoretical imaging contrast enhancement factor between the defect and the background at that point, predicted by the model under a specific polarization state.
[0175] The principle for weight allocation is as follows: the more valuable the information of an image acquired under the polarization state with the higher the model's predicted contrast, the greater its weight should be; and the greater the theoretical contrast difference between two images, the more unique polarization modulation information may be contained in their difference map, and the greater its weight should be.
[0176] Based on this, for each original image I k Assign a base weight , so that it is with Proportional; at the same time, for each difference image D mn Assign a difference weight The absolute value of the difference between it and the corresponding two gain values. Proportional.
[0177] Subsequently, nonlinear normalization is performed to ensure that the sum of all weights is 1, while compressing potentially maximal weight values to prevent them from dominating the fusion result. Specifically, the following formula is used:
[0178] For each original weight w i (including all) and To suppress the influence of excessive weights, a nonlinear compression transformation can be performed, such as taking the square root (or using logarithmic, exponential, or other functions).
[0179] ;
[0180] The weights v after compression transformation i Perform linear normalization to obtain the final weight W. i :
[0181] ;
[0182] Where n is the total number of images involved in the fusion (including the original image and the difference image). This represents the k-th compressed weight value.
[0183] This normalization process compresses the original weights using a nonlinear function (such as the square root), ensuring that large weights are not linearly amplified, thereby improving the robustness of the fusion results.
[0184] Finally, the normalized weight W is obtained. k (corresponding to I) k ) and W mn (Corresponding to D) mn ).
[0185] S403c: Add the weighted original image and the difference image to obtain the value of the pixel in the preliminary fused image F:
[0186] ;
[0187] This computation is performed in parallel on all pixels on the system's image processing unit (such as a GPU).
[0188] By fusing the data, the system not only enhances the defect signal that appears most obvious under the optimal polarization state, but also effectively enhances the edge contour of the defect through differential images. More importantly, the entire process is dynamically driven by the theoretical predictions of the polarization physics model, rather than a fixed routine, thus exhibiting adaptability to different surface geometries and defect types.
[0189] S404: Local contrast enhancement processing.
[0190] After obtaining the initial fused image F, local contrast enhancement based on guided filtering and nonlinear transformation is performed:
[0191] S404a: Use the fused image F itself as a guide image for guided filtering. The advantage of guided filtering is that it smooths the image, estimates a smooth image B with a large range of background illumination components, and at the same time, it can better preserve the edges that are consistent with the structure of the guide image.
[0192] The filter window size and regularization parameters need to be set according to the image noise level and the desired background smoothness. For example, the window radius is usually set between 5 and 15 pixels, and the regularization parameters are selected in a small positive range based on experimental results.
[0193] S404b: After obtaining the smoothed background image B, it is subtracted from the original fused image F to obtain the detail layer image D, which mainly contains high-frequency details, texture and edge information, i.e., D=FB.
[0194] S404c: Adaptive nonlinear enhancement of detail layer D:
[0195] For areas in the detail layer where grayscale changes are gradual and likely belong to a uniform background or noise, apply a small gain to prevent the noise from being excessively amplified.
[0196] For areas with drastic grayscale changes that may correspond to defect edges, a larger gain is applied to significantly stretch their contrast.
[0197] In implementation, a piecewise gain function can be designed for the absolute grayscale value of each pixel in detail layer D. For example, one or more grayscale thresholds can be set:
[0198] When the absolute value of a pixel grayscale is below a certain low threshold, it is considered to be mainly noise and is given a low gain coefficient close to 1 (such as 1.0~1.5).
[0199] When the absolute value of a pixel's grayscale exceeds a certain high threshold, it is considered that it is likely to be a significant feature and is given a high gain coefficient (such as 2.0~3.5).
[0200] In the transition region between these two thresholds, the gain coefficient can increase linearly or non-linearly with the gray value.
[0201] It should be noted that the specific values of the threshold and gain coefficient can be initially determined by statistical analysis of a batch of typical sample images, and fine-tuned during the system debugging phase.
[0202] S404d: The detail layer after nonlinear enhancement is added back to the previously obtained background image B, which maintains overall lighting smoothness, to obtain the final image E with enhanced local contrast.
[0203] S405: Diffraction enhancement for subpixel-level minute defects.
[0204] After the preceding fusion and contrast enhancement, the defect signals in image E have been strengthened. However, for micron-sized scratches and pits that are close to or smaller than the diffraction limit of the optical system, their edges will become blurred in the original image due to light diffraction and unavoidable slight defocus, with information scattered across multiple pixels. Therefore, to further improve the sharpness of these tiny defects, digital refocusing processing based on scalar diffraction theory is introduced.
[0205] Using the already precisely known object surface depth corresponding to each pixel (from the Z-axis of the geometry perception module) ij The optical parameters of the imaging system (such as numerical aperture and wavelength) are also known, and the propagation process of light can be simulated and reversed in the digital domain, thereby solving for the image under ideal focusing conditions.
[0206] S405a: First, based on the depth information Z of each pixel... ij Given the focal plane position of the system, calculate the defocus amount corresponding to that point; then, using the angular spectrum propagation formula in scalar diffraction theory, construct a virtual phase modulation plate function for the entire image.
[0207] This function is essentially a complex matrix that describes the phase delay distribution experienced by light waves as they propagate from the surface of an object to the camera sensor due to defocusing and diffraction. The key inputs to construct this function are the amount of defocus, the wavelength of the light wave, and the aperture function of the system.
[0208] S405b: The enhanced image E is regarded as a light intensity distribution and given an initial uniform phase distribution, which is combined into a complex amplitude field. Then, using the angular spectrum propagation theory, this complex amplitude field is combined with the virtual phase modulation board we constructed in the frequency domain to achieve a digital backpropagation.
[0209] In simple terms, this process attempts to cancel out the wavefront distortion caused by defocus and diffraction in actual imaging. From a signal processing perspective, this is equivalent to a deconvolution operation with a known point spread function, which is determined by the aforementioned optical parameters and the amount of defocus.
[0210] S405c: Since direct deconvolution is sensitive to noise and may be unstable, iterative optimization algorithms are usually used to solve for the most likely ideal focused image.
[0211] Commonly used algorithms, such as the Richardson-Lucy algorithm, are suitable for image restoration problems with known point spread functions. This algorithm starts with the current image estimate and iteratively compares the actual acquired image (after a simulated blurring process) with the image estimated by the algorithm, continuously correcting the estimate until it converges to a mathematically most likely and clearer image.
[0212] The number of iterations can be set according to the requirements for processing speed and restoration accuracy. For example, 10 to 30 iterations can usually achieve significant results.
[0213] Through the aforementioned digital refocusing process, the system can largely restore sub-pixel-level details blurred by optical diffraction and defocusing effects, making the edges of tiny defects sharper and clearer in the image, thereby significantly improving the detection capability and characterization accuracy of defects such as micro-scratches and micro-dimples.
[0214] The defect identification and decision-making module is used to automatically detect and classify defects in the final image after image fusion and enhancement. Its core is a deep convolutional neural network model, supplemented by rule-based post-processing and report generation. The specific implementation process is broken down as follows:
[0215] S501: Structure and reasoning of deep convolutional neural network models.
[0216] This embodiment uses a fully convolutional network with an encoder-decoder structure as the core model to learn the mapping from the input image to the defect probability map in an end-to-end manner.
[0217] S501a: The encoder is responsible for feature extraction and compression, and consists of five cascaded convolutional modules. Each module contains one convolutional layer and one downsampling layer. Specifically, each convolutional layer uses a small 3×3 convolutional kernel and ReLU as the activation function; after convolution, a 2×2 window max pooling layer with a stride of 2 is applied for downsampling.
[0218] To define the capacity of the network, the number of channels (i.e. the number of filters) in the encoder part of the convolutional layer is usually increased layer by layer. For example, it can be set to 64, 128, 256, 512, and 1024 in sequence. Through these 5 downsampling steps, the spatial size of the feature map gradually decreases, while the number of channels and the degree of semantic abstraction gradually increase, so that the network can capture multi-scale feature information from local texture to global structure.
[0219] S501b: The decoder is responsible for restoring the deep, low-resolution feature maps extracted by the encoder to the size of the original input image and outputting a fine segmentation result. The decoder is composed of 5 transposed convolutional layers (also known as deconvolutional layers), which perform upsampling by 2x step by step.
[0220] To prevent the loss of spatial details crucial for segmenting minute defects during multiple downsampling and upsampling processes, the network introduces a skip connection structure. The specific implementation is as follows:
[0221] The high-resolution, detailed feature maps output from each stage of the encoder (before pooling) are directly concatenated in the channel dimension to the feature maps upsampled in the corresponding intermediate stage of the decoder. For example, the output of the first layer of the encoder (64 channels) is concatenated to the feature map upsampled in the last layer of the decoder.
[0222] After stitching, a 1×1 convolutional layer is usually used to fuse and adjust the number of channels to control computational complexity. This allows the decoder to simultaneously utilize the precise location and edge information preserved by the shallow network and the high-level semantic information extracted by the deep network when reconstructing the image size, thereby improving the segmentation accuracy of small, blurred edge defects.
[0223] S501c: The final output layer of the network is a 1×1 convolutional layer followed by a pixel-wise sigmoid activation function. The function of this layer is to map the multi-channel feature map output by the decoder into a single-channel image and compress the value of each pixel to between 0 and 1. Its output is a defect segmentation probability map with the same size as the input image. The value of each pixel in the map intuitively represents the confidence that the position belongs to the defect region.
[0224] S502: Model training and deployment.
[0225] S502a: The above network structure must undergo sufficient offline supervised training before being put into actual online testing.
[0226] The training dataset contains more than 500,000 image samples to ensure the model's generalization ability. The dataset includes both synthetic images generated using computer graphics techniques with precise 3D background shapes and various parametric defects, and real defect images collected in real industrial scenes and meticulously annotated by humans. The combination of the two ensures both the diversity and authenticity of the data.
[0227] The training process aims to minimize the difference between the predicted results and the true labels, using binary cross-entropy as the loss function and Adam (adaptive moment estimation) as the optimization algorithm.
[0228] Key training hyperparameters need to be set to ensure stable convergence. For example, the initial learning rate is usually set between 1e-4 and 1e-3, and a learning rate decay strategy can be used; the batch size is set to 8, 16, or 32 depending on the GPU memory; the number of training iterations (Epochs) usually needs to be hundreds of rounds until the loss function no longer decreases significantly on the validation set.
[0229] S502b: After training, a set of optimal network weight parameters is obtained.
[0230] During system operation, the decision output unit loads these trained model parameter files into memory or dedicated computing hardware (such as GPU). During actual detection, the image output by the image fusion and enhancement module is directly input into the network, which can efficiently calculate the defect probability map of the entire image in parallel during the forward propagation process and complete online inference.
[0231] S503: Decision output process based on probabilistic graphs.
[0232] After the model inferences to obtain the probability map, the decision output unit performs a series of post-processing steps to finally generate a structured detection report.
[0233] S503a: First, set a defect probability threshold (e.g., 0.7, which can be optimized during the model validation phase with the goal of maximizing the overall detection rate and reducing the false alarm rate). Mark all pixels in the probability map with values greater than this threshold as foreground, i.e., potential defect pixels; mark the remaining pixels as background.
[0234] Subsequently, connected component analysis was performed on the binarized foreground region, and a unique label was assigned to each independent connected region, thereby initially determining the number of candidate defects and their pixel-level location in the image.
[0235] S503b: Next, for each marked connected region, calculate its geometric and morphological features, which typically include: converting the pixel area into the actual physical area according to the pixel-physical size conversion relationship predefined by the system; calculating the aspect ratio of the minimum bounding rectangle of the region; calculating the first and second moments (such as Hu moments) based on the region contour to describe the shape characteristics; and statistically analyzing the intensity features such as the average gray value and gray standard deviation within the region.
[0236] S503c: The decision output unit pre-stores a judgment logic based on rules or a simple classifier. This logic classifies each candidate region into several predefined typical defect types based on the feature combinations calculated above. For example:
[0237] If the aspect ratio of a region is significantly greater than a preset threshold (e.g., greater than 3), the physical area is small, and the linearity of the outline is high, then it is classified as a "scratch".
[0238] If the physical area of a region is small, its aspect ratio is close to 1 (e.g., between 0.8 and 1.2), and its outline has a high degree of circularity or elliptic fit, it tends to be classified as a "pit".
[0239] If the area is relatively large, the boundary gradient is gentle, and the internal gray distribution is relatively uniform or has a gradual change characteristic, it may be classified as "stain" or "oil stain".
[0240] The specific classification rules and threshold parameters can be determined through statistical analysis and adjustment based on the historical defect sample database of the actual tested parts.
[0241] S503d: Finally, the module summarizes all detected defect information and generates a standardized structured inspection report.
[0242] Reports are typically recorded in common formats such as XML and JSON. The content includes at least the timestamp of the inspection task, the unique identifier of the inspected part, the overall inspection conclusion (e.g., pass / fail), the total number of defects detected, and a detailed entry for each defect. Each defect entry usually includes the defect type, its three-dimensional spatial coordinates (X, Y, Z) in the part's coordinate system (which can be calculated by combining image coordinates with depth information provided by the geometry module), physical dimensions (e.g., area, length, width), and confidence level (optional).
[0243] Meanwhile, the report can be uploaded in real time to the factory's production execution system or quality management system via a network interface for quality traceability, statistical analysis, and production decision-making.
[0244] In summary, the defect identification and decision-making module achieves pixel-level defect probability prediction through a fully convolutional network, and then combines it with traditional image analysis and rule judgment to complete the final conversion from enhanced image to quantified defect information.
[0245] As a key implementation feature of this embodiment, the entire system operates within a hierarchical closed-loop control architecture, which aims to coordinate the system's ability to respond quickly to instantaneous changes and adapt to long-term drift. Specifically, it includes two control loops operating at different time scales.
[0246] The first is the bottom-level millisecond-level polarization fast response loop, which consists of a geometry sensing module, a dynamic polarization control module, and an optical imaging module connected in series, forming a high-speed closed loop from sensing to execution. Its working cycle is strictly synchronized with the acquisition frame rate of the industrial camera, which is approximately 16.7 milliseconds (corresponding to 60 frames / second) in this embodiment. Within this extremely short cycle, the system sequentially completes the following operations:
[0247] First, the three-dimensional topography and normal information of the part surface are collected and calculated through the geometric topography perception module;
[0248] Subsequently, the dynamic polarization control module calculates and outputs the optimal polarization angle combination command based on this real-time geometric information, driving the polarization device to quickly adjust into position;
[0249] Finally, the optical imaging module acquires the next frame image under the adjusted polarization state.
[0250] This high-speed closed loop ensures that the system can track local geometric changes caused by the movement of the robotic arm or the changes in the surface of the part in real time, thereby achieving adaptive matching of polarization state and ensuring image contrast and clarity.
[0251] The second is the second-level adaptive loop of system parameters located at the upper level. Its operating cycle is relatively long, usually set between 10 and 60 seconds, and it is used to monitor the macroscopic performance of the system.
[0252] The adaptive loop outputs statistical indicators over a period of time from the defect recognition and decision-making module, such as the average contrast of all detected images in the past 30 seconds, the overall defect detection rate, and the false positive alarm rate. These real-time statistical values are compared with preset performance thresholds.
[0253] When a key indicator (such as average contrast) remains below its threshold for more than three monitoring cycles, the adaptive loop is triggered. Its function is not to interfere with the rapid response of the underlying layer, but to initiate a slow, iterative parameter optimization process. The adjustment targets are mainly those relatively stable model or algorithm parameters that may change slowly with the production environment, such as the optical constants of the material in the polarization optimization model (the real part n and the imaginary part k of the complex refractive index), or the weight mapping function parameters of the fusion algorithm in the image fusion and enhancement module.
[0254] The adjustment process employs optimization methods such as gradient descent with very small step sizes. The aim is to allow the overall performance indicators of the system to slowly return to the optimal range, giving the system long-term "learning" and adaptive capabilities. This enables it to effectively cope with the slow, systematic drift caused by factors such as batch replacement of parts, fine-tuning of surface coating processes, and gradual aging of lighting sources over time.
[0255] Through the dual-loop collaboration of millisecond-level fast response and second-level slow adaptation, high-precision, real-time optical control of transient geometric features on the surface of complex curved parts is achieved. It has robustness to adapt to long-term changes in the production environment, fundamentally improving the reliability, stability and maintainability of the online operation of the inspection system. It significantly extends the cycle of frequent manual calibration and parameter adjustment required by traditional vision inspection systems, thereby better meeting the needs of continuous and stable production in industrial sites.
[0256] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, the embodiments should be regarded as exemplary and non-limiting in all respects.
[0257] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment includes only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A visual inspection system for surface defects of mechanical parts, characterized in that, include: An optical imaging module is used to acquire raw image data of the surface of the mechanical part to be inspected; The geometric shape perception module is coaxially integrated with the optical imaging module and is used to calculate the three-dimensional geometric information of the surface of the mechanical part to be inspected in real time. The dynamic polarization control module, connected to the geometric shape perception module and the optical imaging module, is used to dynamically adjust the polarization state of the illumination light incident on the surface of the mechanical part to be inspected and the polarization state of the reflected light received by the camera, based on the three-dimensional geometric information output in real time by the geometric shape perception module. An image fusion and enhancement module, connected to the optical imaging module and the dynamic polarization control module, is used to fuse and enhance the features of at least three multi-frame polarization images acquired under different polarization states controlled by the dynamic polarization control module to obtain an enhanced image. The defect identification and decision-making module, connected to the image fusion and enhancement module, is used to automatically detect and classify defects in the enhanced image and output the detection results; The dynamic polarization control module has a built-in polarization optimization model. The polarization optimization model takes the surface normal vector, incident light direction vector, camera observation direction vector and optical constant of the part material from the geometric shape perception module as input, and takes maximizing the imaging signal contrast between the defect area and the normal background area as the optimization objective. It calculates the optimal incident polarization angle and the optimal receiving polarization angle for each pixel in parallel. The image fusion and enhancement module performs the following operations: Subpixel-level registration of multi-frame polarization images; Calculate the difference images between each pair of registered images; Based on the theoretical contrast gain value calculated and stored for each image pixel by the dynamic polarization control module, adaptive fusion weights are assigned to each original image and difference image and weighted fusion is performed to obtain a preliminary fused image. The preliminary fused image is then subjected to local contrast enhancement processing.
2. The visual inspection system for surface defects of mechanical parts according to claim 1, characterized in that, The optical imaging module includes a high-resolution global shutter industrial camera, a fixed-focus telecentric lens, and a controllable illumination unit. The controllable lighting unit is a multi-layer ring array multi-angle LED light source, which includes at least two coaxial and independently controlled lighting sub-rings. The angle between the normal of the emitting surface of each lighting sub-ring and the optical axis of the system is different, and the brightness can be adjusted independently and continuously.
3. The visual inspection system for surface defects of mechanical parts according to claim 2, characterized in that, The geometric shape perception module merges the structured light projection path and the camera imaging path through a beam splitter. The geometric shape perception module includes a structured light projection unit, which is used to project a set of light stripe patterns encoded in space and time onto the surface of the part. The high-resolution global shutter industrial camera synchronously acquires images of deformed stripes modulated by the three-dimensional shape of the part surface. The geometric shape perception module also includes a depth calculation unit, which processes the deformed stripe image, reconstructs the three-dimensional point cloud of the part surface in real time, and estimates the surface normal direction.
4. The visual inspection system for surface defects of mechanical parts according to claim 3, characterized in that, The dynamic polarization control module includes an incident polarization control unit and a receiving polarization control unit. The incident polarization control unit is disposed at the optical path exit of the controllable illumination unit and includes a linear polarizer and a first driving mechanism for driving its rotation. The receiving polarization control unit is located at the front end of the industrial camera lens and includes an analyzer and a second drive mechanism that drives its rotation.
5. The visual inspection system for surface defects of mechanical parts according to claim 4, characterized in that, The image fusion and enhancement module also performs diffraction enhancement operations for sub-pixel-level minute defects: Based on the depth information of each pixel and the optical parameters of the imaging system provided by the geometric shape perception module, digital refocusing is performed using scalar diffraction theory to remove the blurred details caused by complex optical diffraction and defocusing effects.
6. The visual inspection system for surface defects of mechanical parts according to claim 1, characterized in that, The defect identification and decision module includes a fully convolutional network model with an encoder-decoder structure, used to perform pixel-level defect probability prediction on the enhanced image and generate a defect segmentation probability map. Based on the defect segmentation probability map, binarization, connected component analysis, feature extraction, and classification are performed to generate a structured inspection report.
7. The visual inspection system for surface defects of mechanical parts according to claim 6, characterized in that, The geometric shape perception module, dynamic polarization control module and optical imaging module are connected in series to form a millisecond-level polarization fast response closed loop, whose working cycle is synchronized with the acquisition frame rate of the high-resolution global shutter industrial camera. The output of the defect identification and decision-making module is connected to a second-level system parameter adaptive loop, which is used to optimize the parameters in the polarization optimization model or the image fusion and enhancement module based on historical detection performance statistics.
8. The visual inspection system for surface defects of mechanical parts according to claim 7, characterized in that, The dynamic polarization control module is also used for: Cluster analysis is performed on the surface normal vectors of each pixel in the current image output by the geometric shape perception module. Pixels with similar normal directions are divided into the same control region. The optimal incident polarization angle and the optimal received polarization angle are calculated for each control region. Then, the optimal polarization angle combination command for each region is generated and sent to the incident polarization control unit and the received polarization control unit for execution.
Citation Information
Patent Citations
Machine vision dynamic defect detection method and device for precise structural part
CN119887745A
Quartz stone surface defect detection method, system and equipment
CN120404749A