Harness connector sealing detection system based on machine vision and deep learning

CN122545002APending Publication Date: 2026-08-11DONGGUAN HAOXIN AUTOMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-11
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0006]针对现有技术的不足,本发明提供了基于机器视觉和深度学习的线束连接器密封性检测系统,解决了现有检测方案难以将盲孔内部的光学离焦模糊与密封栓的真实物理形变相分离,单凭二维特征容易导致密封性判定准确率低的问题

Benefits of technology

1、本发明通过变焦扫描获取原始焦栈图像序列,利用局部窗口方差与高斯拟合生成拓扑深度图和全聚焦纹理图。该机器视觉提取方式克服了离散焦距步长带来的分辨率限制,重构出线束连接器内部密封栓的亚像素级深度与形貌特征,解决了盲孔环境下图像对比度低导致特征提取困难的问题,为密封性检测提供了可靠的数据基础。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122545002A_ABST
    Figure CN122545002A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of image processing, and discloses a wire harness connector sealing property detection system based on machine vision and deep learning, which comprises a data acquisition module, an original focal stack image sequence of a wire harness connector is acquired through zoom scanning; a depth reconstruction module, after registration of the original sequence, local window variance is calculated and Gaussian fitting is performed to generate a topological depth map, an out-of-focus gradient parameter and a full-focus texture map; a strain calculation module, high-frequency energy features of the full-focus texture map are corrected by using the out-of-focus gradient parameter to obtain a spatial frequency, and a microscopic strain field tensor map is generated; and a defect classification module, the microscopic strain field tensor map is converted into an attention weight matrix, the attention weight matrix is weightedly fused with an extracted geometric feature map, and a detection result is mapped and output. The application decouples optical defocusing and physical deformation by using machine vision, uses strain prior to guide deep learning fusion, overcomes the limitation of two-dimensional feature judgment, and realizes accurate detection of the sealing property of the wire harness connector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, specifically to a wire harness connector sealing performance detection system based on machine vision and deep learning. Background Technology

[0002] The assembly condition of the internal rubber sealing plugs in wire harness connectors determines the overall sealing performance. Sealing performance testing is used to identify defects such as insufficient sealing plug assembly depth or edge curling. Machine vision, which uses industrial cameras to acquire images and perform calculations, combined with deep learning, which uses neural networks to extract features for classification, is currently the main technology for automated inspection.

[0003] Existing sealing performance testing solutions typically use fixed-focal-length machine vision equipment to acquire images of the connector's interior. These images are then input into a deep learning network, which extracts the surface texture and geometric contour features of the sealing bolts on a two-dimensional plane to directly assess the current assembly status and output sealing performance results.

[0004] Because wire harness connectors have a deep, narrow, blind-hole structure, the depth of field of the lens limits the amount of time, causing localized optical defocus blur in the captured images. When the sealing plug is compressed and undergoes real physical deformation, the surface texture changes and optical blur overlap in the image, making it difficult for existing technologies to separate optical image attenuation from pure physical deformation. Furthermore, existing deep learning networks lack data guidance on the physical stress state, making it difficult to accurately locate the deformation area in dark environments relying solely on two-dimensional appearance. Deep learning networks are prone to misinterpreting feature loss caused by defocus as physical defects, resulting in low accuracy in wire harness connector sealing detection.

[0005] Therefore, this invention proposes a wire harness connector sealing performance detection system based on machine vision and deep learning to address the shortcomings of existing technologies. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention provides a wire harness connector sealing performance testing system based on machine vision and deep learning. This system solves the problem that existing testing methods struggle to separate the optical defocusing blur inside blind holes from the actual physical deformation of the sealing plug, and that relying solely on two-dimensional features easily leads to low accuracy in sealing performance determination.

[0007] To achieve the above objectives, the present invention provides the following technical solution: a wire harness connector sealing performance testing system based on machine vision and deep learning, comprising: The data acquisition module is used to acquire the original focal stack image sequence of the wire harness connector through zoom scanning; The depth reconstruction module is used to perform phase registration on the original focal stack image sequence to obtain an aligned focal stack sequence, calculate the local window variance of the aligned focal stack sequence and perform Gaussian fitting to generate a topological depth map, defocus gradient parameters and full-focus texture map. The strain calculation module is used to calculate the high-frequency energy characteristics of the full-focus texture map, use the defocus gradient parameter to correct the high-frequency energy characteristics to obtain the normalized spatial frequency, and generate a micro-strain field tensor map based on the normalized spatial frequency. The defect classification module is used to extract the geometric feature maps of the topological depth map and the full-focus texture map, convert the micro-strain field tensor map into an attention weight matrix and perform weighted fusion with the geometric feature map to generate a fused feature map, and map the fused feature map to output the sealing detection result.

[0008] Preferably, the data acquisition module is configured with an orthogonal polarization coaxial light source, a programmable zoom lens, and an industrial camera; the industrial camera drives the programmable zoom lens to continuously adjust the focal length according to a preset focal length step, and triggers a global shutter on each set focal length plane to perform exposure imaging and acquire two-dimensional grayscale images; the data acquisition module aligns and superimposes the continuously captured two-dimensional grayscale images along spatial coordinates with the preset focal length step to construct the original focal stack image sequence.

[0009] Preferably, the depth reconstruction module extracts the intermediate layer image from the original focal stack image sequence as the reference frame image, performs a two-dimensional discrete Fourier transform on the reference frame image and each frame to be registered in the original focal stack image sequence to obtain the reference frequency domain complex matrix and the to-be-registered frequency domain complex matrix, respectively, calculates the cross power spectrum of the reference frequency domain complex matrix and the to-be-registered frequency domain complex matrix and performs an inverse Fourier transform to obtain the impulse response function in the spatial domain, extracts the two-dimensional coordinates when the impulse response function in the spatial domain reaches an extreme value as the translation compensation vector, and uses a bilinear interpolation algorithm to perform spatial interpolation translation on each frame to be registered according to the translation compensation vector to output the aligned focal stack sequence.

[0010] Preferably, the depth reconstruction module uses a Laplacian convolution kernel to perform a convolution operation on each discrete focal length layer in the aligned focal stack sequence to obtain a gradient image. A local neighborhood window is set on the gradient image, and the variance of the gradient values ​​within the local neighborhood window is calculated as the local window variance. The focal length index corresponding to the maximum value of the local window variance is extracted along the optical axis direction. The focal length index is mapped to physical space depth by combining a preset focal length stride to generate the topological depth map.

[0011] Preferably, the depth reconstruction module performs one-dimensional Gaussian fitting on the curve formed by the local window variance along the optical axis to transform it into a continuous Gaussian curve function, extracts the symmetry axis coordinates of the continuous Gaussian curve function, multiplies them by the preset focal length step size to update the topological depth map, extracts the standard deviation of the continuous Gaussian curve function as the defocus gradient parameter, and simultaneously extracts the gray value corresponding to the symmetry axis coordinate by linear interpolation of the corresponding pixels of adjacent discrete layer images along the optical axis, and performs pixel stitching in the spatial dimension to generate the fully focused texture map.

[0012] Preferably, the strain calculation module divides the fully focused texture map into multiple non-overlapping grid blocks in a spatial dimension, performs a two-dimensional fast Fourier transform on each grid block, shifts the zero-frequency component to the center of the spectrum to obtain the frequency domain distribution matrix, calculates the spectral energy outside the set cutoff frequency radius, and generates the high-frequency energy features.

[0013] Preferably, the strain calculation module retrieves the defocus gradient parameters in two-dimensional matrix form corresponding to the grid block to calculate the local defocus amount, uses the local defocus amount to perform division compensation calculation on the high-frequency energy characteristics to obtain the normalized spatial frequency, retrieves the pre-stored reference frequency distribution matrix to calculate the relative rate of change between the normalized spatial frequency and the reference frequency distribution matrix to generate a relative tensor value, and arranges the relative tensor values ​​of each grid block into a two-dimensional matrix according to the two-dimensional grid coordinate sequence to generate the microscopic strain field tensor map.

[0014] Preferably, the defect classification module performs size normalization preprocessing on the topological depth map and the fully focused texture map, and directly splices them in the channel dimension to generate an initial bimodal tensor. The initial bimodal tensor is input into the shallow feature extractor of the cross-modal feature fusion convolutional neural network, and the local texture of the initial bimodal tensor is extracted through local receptive field sliding window operation to output the geometric feature map.

[0015] Preferably, the defect classification module adjusts the spatial resolution of the micro-strain field tensor map using bilinear interpolation upsampling, inputs the spatially adjusted micro-strain field tensor map into the physical prior mapping network branch, generates the attention weight matrix through cross-channel integration and nonlinear activation function compression transformation, and generates the fused feature map by performing element-wise dot product operation on the geometric feature map and the attention weight matrix in the spatial dimension at the feature level.

[0016] Preferably, the defect classification module inputs the fused feature map into a deep classification head network for feature dimensionality reduction and abstraction, compresses it into a one-dimensional dense vector through a global average pooling layer, and sends it into a fully connected layer to calculate the logical value of each category. The vector is then converted into a classification probability distribution vector through an activation function, and the index mapping corresponding to the highest probability value in the classification probability distribution vector is extracted to output the sealing detection result.

[0017] This invention provides a wire harness connector sealing performance testing system based on machine vision and deep learning. It offers the following advantages: 1. This invention acquires the original focal stack image sequence through zoom scanning and generates a topological depth map and a full-focus texture map using local window variance and Gaussian fitting. This machine vision extraction method overcomes the resolution limitations caused by discrete focal length steps, reconstructs the sub-pixel level depth and morphological features of the sealing bolts inside the wire harness connector, and solves the problem of low image contrast leading to difficulty in feature extraction in blind hole environments, providing a reliable data foundation for sealing performance detection.

[0018] 2. This invention extracts high-frequency energy features from the full-focus texture map and uses defocus gradient parameters obtained from machine vision reconstruction for compensation and correction, generating a microscopic strain field tensor map. This decoupled calculation method eliminates optical blur interference caused by lens depth-of-field changes, quantitatively reflects the true deformation state of the sealing bolt during assembly, solves the problem that conventional inspection easily misjudges optical blur as physical sealing defects, and improves the authenticity of defect features in wire harness connectors.

[0019] 3. This invention constructs a deep learning feature fusion network to weight and fuse visually extracted geometric feature maps with attention weight matrices transformed from microscopic strain field tensor maps. This deep learning judgment method utilizes prior physical stress data to guide the network to focus on spatial regions with abnormal deformation, overcoming the limitations of relying solely on two-dimensional appearance features for judgment, and achieving objective classification and output of sealing defects in wire harness connectors. Attached Figure Description

[0020] Figure 1 This is an architecture diagram of the wire harness connector sealing performance detection system based on machine vision and deep learning according to the present invention. Figure 2 This is a flowchart of the wire harness connector sealing performance detection method based on machine vision and deep learning according to the present invention. Figure 3 This is a graph showing the variance of a single pixel's local window and the Gaussian fitting curve according to the present invention.

[0021] Among them, 100 is the data acquisition module; 200 is the deep reconstruction module; 300 is the strain calculation module; and 400 is the defect classification module. Detailed Implementation

[0022] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] See attached document Figure 1 This invention provides a wire harness connector sealing performance testing system based on machine vision and deep learning, the testing system comprising: The data acquisition module 100 is used to acquire the original focal stack image sequence of the wire harness connector through zoom scanning; The depth reconstruction module 200 is used to perform phase registration on the original focal stack image sequence to obtain the aligned focal stack sequence, calculate the local window variance of the aligned focal stack sequence and perform Gaussian fitting to generate a topological depth map, defocus gradient parameters and full-focus texture map. The strain calculation module 300 is used to calculate the high-frequency energy features of the full-focus texture map, and to correct the high-frequency energy features using the defocus gradient parameter to obtain the normalized spatial frequency, and to generate a micro-strain field tensor map based on the normalized spatial frequency. The defect classification module 400 is used to extract geometric feature maps from the topological depth map and the full-focus texture map, convert the micro-strain field tensor map into an attention weight matrix and perform weighted fusion with the geometric feature map to generate a fused feature map, and map the fused feature map to output the sealing detection result.

[0024] See attached document Figure 2 This invention provides a method for detecting the sealing performance of wire harness connectors based on machine vision and deep learning, comprising the following steps: S10, acquire the original focal stack image sequence of the wire harness connector through zoom scanning; S20: Phase registration is performed on the original focal stack image sequence to obtain the aligned focal stack sequence. The local window variance of the aligned focal stack sequence is calculated and Gaussian fitting is performed to generate the topological depth map, defocus gradient parameters and full focus texture map. S30: Calculate the high-frequency energy features of the full-focus texture map, use the defocus gradient parameter to correct the high-frequency energy features to obtain the normalized spatial frequency, and generate a micro-strain field tensor map based on the normalized spatial frequency. S40: Extract the geometric feature maps of the topological depth map and the full-focus texture map, convert the micro-strain field tensor map into an attention weight matrix and perform weighted fusion with the geometric feature map to generate a fused feature map, and map the fused feature map to output the sealing detection result.

[0025] The wire harness connector sealing performance detection method and system based on machine vision and deep learning of this invention belong to the same inventive concept. Each logical module in the detection system is configured to execute the corresponding steps in the detection method. Specifically, the data acquisition module 100 in the detection system is used to execute step S10 to acquire a three-dimensional original focus stack image sequence; the depth reconstruction module 200 is used to execute step S20 to extract spatial depth and obtain defocus gradient parameters; the strain calculation module 300 is used to execute step S30 to complete the mathematical decoupling of optical defocus and physical deformation features; and the defect classification module 400 is used to execute step S40 to complete cross-modal feature weighted fusion and classification. The detection system relies on the sequential interactive operation of each module to achieve a computational closed loop from optical data acquisition and physical parameter decoupling to sealing defect identification.

[0026] To further clarify the implementation of each technical aspect of the present invention, the following will provide a detailed description of the implementation of each functional module involved above and its internal processing flow.

[0027] See attached document Figure 1 and Figure 2 In this embodiment, to achieve effective feature extraction in low-contrast environments, the data acquisition module 100 is configured with an orthogonally polarized coaxial light source, a programmable zoom lens, and an industrial camera at the physical imaging level. The orthogonally polarized coaxial light source, the programmable zoom lens, and the industrial camera can all utilize existing mature commercial hardware devices in the field. As a preferred embodiment, a polarizer is arranged at the emitting end of the orthogonally polarized coaxial light source, and an analyzer is arranged at the front end of the lens of the industrial camera. The polarization direction of the polarizer and the polarization direction of the analyzer are perpendicular to each other, forming an orthogonal polarization state.

[0028] In the wire harness connector inspection station, the plastic hole wall material of the wire harness connector maintains its original polarization state when reflecting light. Because the polarizer and analyzer are orthogonally positioned, the direct specular reflection light maintaining its original polarization state is physically blocked by the analyzer, filtering out background interference from ambient highlights and the inner wall structure of the blind hole. The surface of the rubber sealing plug inside the wire harness connector has a micro-undulating texture structure. This micro-undulating texture structure causes diffuse reflection of the incident polarized light and induces a depolarization effect. The depolarized light contains polarization components in multiple directions, which can penetrate the analyzer and be captured by the target surface of the industrial camera. By configuring orthogonal polarization light paths, the data acquisition module 100 extracts the high-frequency texture details of the rubber sealing plug inside the hole, providing high signal-to-noise ratio basic image data for computational processing. Regarding the specific installation structure and fixing bracket design of the polarizer and analyzer, those skilled in the art can make conventional selections and mechanical adaptations according to the actual workstation space; these are well-known technologies in the field and will not be elaborated further here.

[0029] After the optical acquisition hardware structure is completed, the data acquisition module 100 performs a zoom scanning operation, which is implemented in the following steps: S101, the industrial camera sends a control electrical signal to the programmable zoom lens, driving the focal plane of the programmable zoom lens to jump to a preset initial scanning position. In this embodiment, the preset initial scanning position is determined by physical height measurement of the blind hole opening of the wire harness connector, for example, set at 2 mm above the upper edge of the blind hole, to ensure that the scanning range completely covers the highest possible position of the sealing plug, avoiding data loss dead zones caused by the micro-flanging of the sealing plug exceeding the hole opening. The programmable zoom lens is specifically a liquid lens based on the electrowetting effect, which adjusts the curvature of the interface between two liquids inside the lens according to the change of applied voltage to achieve switching of focal length without mechanical movement.

[0030] S102, an industrial camera, operates according to a preset focal length step. Within a set time window, the programmable zoom lens is continuously adjusted along the optical axis. Within a set scanning depth range... The data acquisition module 100 needs to acquire a total number of image frames. Determined according to the following formula: ; In the formula, This indicates the total number of image frames acquired; This indicates a round-down operation; This indicates the set scanning depth range, which is the pre-calibrated physical depth scanning range of the blind hole of the wire harness connector. It is determined based on the internal design and installation tolerances of the wire harness connector, for example, it is set to 15 mm. This represents the preset focal length step, which is the physical span of focal length between two adjacent exposures. Its specific value is determined by the depth-of-field range of the programmable zoom lens, typically set to half the optical depth of field, for example, between 0.1 mm and 0.2 mm. This ensures that there is an overlapping sharp area between adjacent frames, maintaining the continuity of data acquisition. Through this calculation logic, the detection system can dynamically adapt to the size specifications of different wire harness connector models, accurately allocating the total number of optical scan frames.

[0031] S103, the industrial camera triggers a global shutter at each set focal length plane for exposure imaging, acquiring a two-dimensional grayscale image at the corresponding depth of focus. Specifically, the focal length plane refers to the two-dimensional spatial cross-section where the programmable zoom lens achieves the clearest optical focus under a specific driving voltage. As the zoom scan progresses, this cross-section shifts sequentially along the optical axis of the blind hole in the wire harness connector. When the lens focuses on the focal length plane corresponding to each set step, the data acquisition module 100 sends a synchronization electrical signal to the industrial camera to trigger the global shutter. The global shutter controls all photosensitive pixel arrays of the image sensor inside the industrial camera to start and end exposure simultaneously at the same absolute moment. By controlling the synchronous photosensitive and electrical signal readout of all pixels, the industrial camera effectively avoids line-by-line spatial distortion and motion blur caused by low-frequency mechanical vibrations on site, thereby accurately outputting a two-dimensional grayscale image that reflects the true optical distribution of the current depth of focus cross-section. The data acquisition module 100 will continuously capture the images... Two-dimensional grayscale images are aligned and superimposed along spatial coordinates and focal length steps to construct a three-dimensional original focal stack image sequence. .in, and Represents the horizontal and vertical coordinates of pixels in a two-dimensional grayscale image. This indicates the corresponding focal length step number. Original focal stack image sequence. Each pixel in the data is discretized to record the degree of optical blur and depth information along the optical axis, thus completing the data acquisition of the structural features of the wire harness connector.

[0032] In real-world industrial environments, mechanical vibrations can easily cause slight spatial shifts between different layers in the original focus stack image sequence. If depth fusion is performed directly, these shifts can lead to ghosting at texture edges and errors in depth calculation. (See attached image.) Figure 1 and Figure 2 Therefore, the processing logic inside the deep reconstruction module 200 in this embodiment is specifically detailed as follows: S201 performs frequency domain phase registration based on cross-power spectrum. The deep reconstruction module 200 extracts the original focal stack image sequence. The intermediate layer image in the image is used as the reference frame image, and the reference frame image is denoted as... ,in This indicates the corresponding focal length step number. The depth reconstruction module pairs each frame to be registered from the reference frame image and the original focal stack image sequence. Performing a two-dimensional discrete Fourier transform converts the image from the spatial domain to the frequency domain, yielding the corresponding frequency domain complex matrices. and .in, and This represents the horizontal and vertical frequency coordinates in the frequency domain. Based on this, the spatial offset between images is evaluated by calculating their cross-power spectrum. Cross-power spectrum The calculation formula is as follows: ; In the formula, Represents a complex matrix in the frequency domain Complex conjugate; absolute value symbol This indicates the magnitude of the matrix elements being calculated; This represents a very small positive number set to prevent the denominator from being zero, as a preferred method. The value is 10 -5 This ensures the algorithm's operational stability in flat regions. The deep reconstruction module contains 200 pairs of cross-power spectra. Perform an inverse Fourier transform to obtain the impulse response function in the spatial domain, and extract the two-dimensional coordinates (i.e., the coordinate offset relative to the origin of the image center) of the impulse response function when it reaches an extremum as the translation compensation vector. After obtaining the translation compensation vectors for all frames, the depth reconstruction module 200 uses bilinear interpolation to perform spatial interpolation translation on each frame of the original focal stack image sequence, outputting a spatially aligned focal stack sequence. The aforementioned extreme value extraction and bilinear interpolation algorithms based on cross-power spectrum belong to existing technologies in the field of image registration. This invention utilizes the high sensitivity of the aforementioned existing technologies to translation characteristics in the frequency domain to effectively decouple and eliminate interlayer rigid offset errors introduced by low-frequency mechanical vibrations in industrial settings, providing a reliable alignment benchmark for subsequent sub-pixel-level depth reconstruction.

[0033] S202, local window variance calculation and depth mapping are performed. To evaluate the focus level of pixels, the depth reconstruction module 200 uses the Laplacian operator to extract high-frequency edge responses of the aligned focus stack sequence. Specifically, for the aligned focus stack sequence... Each discrete focal length The depth reconstruction module 200 uses a Laplacian convolution kernel to perform a convolution operation on the image to obtain a gradient image. As a preferred method, the Laplacian convolution kernel uses a standard 3×3 matrix, specifically defined as a matrix with a center element of -4, adjacent elements of 1, and diagonal elements of 0, to extract the second derivative information of the image. Subsequently, the gradient image is plotted using coordinates... Define a local neighborhood window with a fixed side length centered on the pixel, and calculate the variance of the gradient values ​​within this local neighborhood window, denoted as the local window variance of the pixel at the current focal length. As a preferred approach, the side length of the local neighborhood window is set to 5 or 7 pixels to balance computational efficiency and noise resistance.

[0034] After the calculation is completed, each layer is traversed along the optical axis, and the focal length index corresponding to the maximum value of the local window variance is extracted. If multiple identical maximum values ​​exist, the first appearing focal length index is extracted by default to avoid the algorithm entering a dead zone. The depth reconstruction module 200 combines the extracted focal length index with a preset focal length step size. Mapped to physical spatial depth, a topological depth map describing the internal three-dimensional morphology of the wire harness connector is generated. The mapping method of the topological depth map is expressed as follows: ; S203 performs Gaussian fitting and feature deconstruction. Because lens imaging has a certain depth of field, the variance of the same pixel within adjacent focal length layers exhibits a non-linear, peak-like distribution. To overcome the vertical resolution limitation imposed by the discrete focal length step size, the depth reconstruction module 200 performs one-dimensional Gaussian fitting on the curve formed by the local window variances along the focal length axis in the aligned focal stack sequence, transforming it into a continuous Gaussian curve function: ; In the formula, ; Represented by natural constant The natural exponential function with base 0; The peak amplitude coefficient represents the value of the fitted Gaussian curve. Represents the coordinates of the axis of symmetry of the fitted Gaussian curve; This represents the standard deviation of the fitted Gaussian curve; This represents the constant bias term generated by the ambient noise at the substrate.

[0035] In this fitted structure, the coordinates of the axis of symmetry This characterizes the precise focus position at the sub-pixel level. The depth reconstruction module 200 further utilizes this symmetry axis coordinate. Multiply by focal length step The topological depth map is updated to output a high-precision sub-pixel depth distribution. Simultaneously, the depth reconstruction module 200 performs linear interpolation along the optical axis on corresponding pixels of adjacent discrete layer images, extracting the grayscale value corresponding to the Gaussian curve's symmetry axis, and then stitches the pixels together in spatial dimension to generate a fully focused texture map. This fully focused texture map integrates sharp pixels of different depths onto the same two-dimensional plane, eliminating local blurring caused by optical defocus.

[0036] More importantly, the standard deviation of the fitted Gaussian curve Physically, this reflects the rate of resolution decay of a pixel along the optical axis, i.e., optical depth-of-field latitude. The depth reconstruction module 200 extracts the standard deviation of the Gaussian curve as the defocus gradient parameter and stores it as a two-dimensional matrix. The defocus gradient parameter not only reflects the optical characteristics of the device but also includes information on beam truncation caused by the occlusion of the inner wall of the aperture. By acquiring and outputting the defocus gradient parameter, the depth reconstruction module 200 provides a quantitative data foundation for subsequent strain calculations to eliminate optical interference and separate purely physical strain features. For the image bilinear interpolation algorithm, those skilled in the art can make conventional selections based on available computing resources; this is a well-known technique in the field and will not be elaborated upon here.

[0037] The rubber sealing pins inside the wire harness connector are physically compressed during insertion and assembly. The microscopic textures on their surface become denser due to deformation, resulting in a spatial frequency shift towards higher frequencies in the frequency domain. In actual optical imaging, due to the depth-of-field attenuation characteristics of the lens, optical defocus also leads to the loss of high-frequency components in the image. To extract the pure physical deformation state, refer to the appendix... Figure 1 and Figure 2 In this embodiment, the strain calculation module 300 establishes a mathematical decoupling mechanism between defocusing and strain characteristics. The specific implementation process is detailed in the following steps: S301, Frequency domain extraction of high-frequency energy features of texture. The strain calculation module 300 receives the fully focused texture map output by the depth reconstruction module 200 and divides the fully focused texture map into multiple non-overlapping grid blocks in a proportional spatial dimension. As a preferred method, the size of the grid blocks is set to 16×16 pixels or 32×32 pixels to achieve a balance between spatial resolution and frequency domain calculation accuracy. For the divided... grid blocks The strain calculation module 300 performs a two-dimensional fast Fourier transform on it and shifts the zero-frequency component to the center of the spectrum to obtain the corresponding frequency domain distribution matrix. Based on this, high-frequency energy characteristics are generated by calculating the spectral energy outside the cutoff frequency. The calculation formula is as follows: ; In the formula, and Represents two-dimensional frequency domain coordinates with the center of the spectrum as the origin; This indicates the set cutoff frequency radius. Cutoff frequency radius The value is determined based on the basic texture period under stress-free conditions of the sealing bolt, and is usually set to one-third of the maximum radius of the frequency domain matrix. (Absolute value symbol) This represents the modulus of the complex number in the frequency domain. For the specific algorithm implementation and calculation process of the two-dimensional Fast Fourier Transform, those skilled in the art can directly call existing digital signal processing library functions for calculation; this is well-known technology in the field and will not be elaborated upon here.

[0038] S302, decoupling correction calculation of defocus and strain characteristics. The strain calculation module 300 synchronously retrieves the value corresponding to the... The defocus gradient parameters of each grid block are represented by a two-dimensional matrix. The average value of the parameters within the grid is calculated and denoted as the local defocus amount. Based on the high-frequency attenuation law of Gaussian blur on the modulation transfer function, the strain calculation module 300 constructs a nonlinear correction model, utilizing local defocusing amount. High-frequency energy characteristics By performing division compensation calculations, the normalized spatial frequency that eliminates optical interference is obtained. : ; In the formula, Represented by natural constant The natural exponential function with base 0; Indicates the calibration coefficient characterizing the optical attenuation properties; This represents a constant bias term set to prevent minimization of the denominator. As a preferred method, the calibration coefficient α is set based on the attenuation curvature of the optical lens's point spread function in the high-frequency band. It is related to the lens's numerical aperture and the observation wavelength, and in this embodiment, it is set between 0.5 and 1.5. Its optimal value can be determined by pre-shooting a standard resolution test target at a known depth out of focus and fitting the high-frequency attenuation curve using the least squares method; constant bias term. Set to 10 -4 Through the above calculation formula, the detection system can adaptively amplify the residual high-frequency energy in the defocused area where optical blur is severe, thereby achieving effective decoupling between physical deformation characteristics and optical imaging interference.

[0039] S303, Differential generation of the micro-strain field tensor map. This involves obtaining the normalized spatial frequencies of all grid blocks. Then, the strain calculation module 300 retrieves the reference frequency distribution matrix pre-stored in the detection system. Reference frequency distribution matrix As constant reference data, it is obtained by performing the same mesh generation and high-frequency energy extraction process as described above on pre-acquired images of standard, undeformed, qualified rubber sealing bolt samples. This data serves as the baseline high-frequency energy value representing the corresponding mesh location under ideal conditions. The strain calculation module 300 calculates the relative rate of change between the two, generating a relative tensor value representing the degree of physical deformation. : ; After the calculation is completed, the strain calculation module 300 calculates the relative tensor values ​​of each grid block. A two-dimensional matrix is ​​arranged according to the two-dimensional grid coordinate sequence of each grid block in the full-focus texture map to generate a micro-strain field tensor map. In this micro-strain field tensor map, the positive value region reflects the high-frequency texture caused by the compression and contraction of the sealing bolt at that location, while the negative value region reflects the low-frequency texture caused by abnormal stretching of the structure. The micro-strain field tensor map output by the strain calculation module 300 transforms the abstract frequency domain energy into a stress distribution spectrum with physical spatial correspondence, providing intuitive and quantitative prior physical data support for subsequent classification of sealing defects.

[0040] See attached document Figure 1 and Figure 2 In this embodiment, to address the challenge of complex morphology and extremely low contrast in simple optical images of sealing defects in wire harness connectors, the defect classification module 400 constructs a cross-modal feature fusion convolutional neural network. This cross-modal feature fusion convolutional neural network jointly solves optical geometric features and physical strain features. Its specific internal network hierarchy, data flow, and model training process are detailed in the following steps: S401, perform network extraction of shallow geometric feature maps. The defect classification module 400 performs unified size normalization preprocessing on the topology depth map and the full-focus texture map output by the depth reconstruction module 200. As a preferred method, the spatial resolution of both is adjusted to 256×256 pixels using bilinear interpolation to meet the computational requirements of the fixed input dimension of the neural network. After size adjustment, the defect classification module 400 directly concatenates the single-channel topology depth map and the single-channel full-focus texture map along the channel dimension to generate an initial bimodal tensor with two input channels. The initial bimodal tensor is then input into the shallow feature extractor of the cross-modal feature fusion convolutional neural network. The shallow feature extractor contains a series of 3×3 standard convolutional layers, batch normalization layers, and linear rectified activation functions, as well as a max pooling layer for size downsampling (the above-mentioned series-connected structure is a common standard convolutional module in this field, such as the basic feature extraction unit widely used in ResNet or VGG networks, which will not be described in detail here). By employing a sliding window operation with multi-layer convolutional kernels for the local receptive field, a shallow feature extractor extracts local textures representing the outer contour of the sealing bolt and the varying depths within the hole, outputting a geometric feature map with a spatial resolution downsampled to 128×128 and a channel dimension expanded to 64. .

[0041] S402, Constructing an Independent Branch Mapping of Physical Prior Weights. To guide the neural network to focus its computation on structural regions with physical stress anomalies, the defect classification module 400 constructs a physical prior mapping network branch independent of the aforementioned shallow feature extractor. The microscopic strain field tensor map output by the strain calculation module 300 is also adjusted to a spatial resolution of 128×128 pixels using bilinear interpolation upsampling and input into this physical prior mapping network branch. The physical prior mapping network branch contains only one set of 1×1 convolutional layers and a Sigmoid nonlinear activation function. Utilizing the cross-channel integration capability of the 1×1 convolutional layers and the squeezing characteristic of the Sigmoid function, the relative tensor values ​​contained in the microscopic strain field tensor map are smoothly compressed and transformed into the real number interval of 0 to 1, generating an attention weight matrix with a dimension of 128×128×1. In this attention weight matrix, the high-weight pixel coordinates with values ​​close to 1 precisely correspond to texture blocks that have experienced excessive physical deformation, such as abnormal compression or overstretching. Since this embodiment employs a spatial domain-based attention mechanism, this attention weight matrix is ​​equivalent to the mapping output of a single spatial attention head.

[0042] S403, completes cross-modal attention fusion and network classification. After acquiring the above features, the defect classification module 400 performs a weighted fusion mechanism at the feature level, combining the geometric feature maps... With attention weight matrix Perform element-wise dot product operations in the spatial dimension: ; In the formula, This represents the generated fused feature map; This represents the corresponding two-dimensional pixel coordinates; The channel index represents the geometric feature map. Through the matrix-dot product operation described above, the defect classification module 400 suppresses the geometric background features of the normal stress region and amplifies the feature response amplitude of the potential defect region.

[0043] After the fusion is complete, the defect classification module 400 will fuse the feature maps. The input is fed into a deep classification head network for feature dimensionality reduction and abstraction. In one embodiment, this deep classification head network consists of four residual blocks connected in series, with the hidden layer channel dimensions set to 128, 256, 512, and 1024 respectively. The abstracted features are sequentially compressed into a one-dimensional dense vector of dimension 1024 through a global average pooling layer, and then fed into a fully connected layer with four nodes to calculate the logical values ​​of each category. Finally, the vector is transformed into a classification probability distribution vector through a Softmax function. The defect classification module 400 uses the Argmax function to extract the index corresponding to the highest probability value in the output vector and outputs the corresponding discretized sealing performance detection result. The sealing performance detection result output in conjunction with the actual business scenario is specifically mapped to four physical states: normal sealing state (meaning that the position and surface deformation of the sealing plug are within the design tolerance), insufficient depth state (meaning that the sealing plug has not been pushed into the specified depth and there is a risk of falling off), micro-flanging state (meaning that the edge of the sealing plug is curled and stuck in the inner wall of the blind hole), and internal twisting state (meaning that the main body of the sealing plug undergoes asymmetrical deformation). Furthermore, depending on the changes in the types of defective products on the actual production line, the aforementioned physical states can be expanded to include other anomalies such as missing sealing plugs and foreign object intrusion, based on actual business needs. For the jump connection structure within the residual block and the specific matrix pooling mechanism of global average pooling, those skilled in the art can refer to the standard ResNet architecture for conventional configuration; these are well-known technologies in the field and will not be elaborated upon here.

[0044] S404, Perform the end-to-end defect classification model training step. Before actual industrial deployment, the detection system needs to iteratively optimize the weight parameters (i.e., the learnable numerical matrices contained in each convolutional kernel and fully connected layer within the network) of the aforementioned cross-modal feature fusion convolutional neural network using the training dataset. This iterative optimization process is essentially a process of using precisely labeled sample data as supervision signals to continuously calculate errors and correct the network weight matrix. As a preferred approach, the initial weights of the shallow feature extractor and the deep classification head network are initialized using weight parameters pre-trained on publicly available industrial defect datasets (such as the MVTec AD dataset) to accelerate model convergence.

[0045] To obtain reliable standard data for error calculation, the training dataset consists of real wire harness connector blind hole image data collected during the production line trial operation phase. All images were preprocessed to generate corresponding topology depth maps, full-focus texture maps, and microscopic strain field tensor maps. Technicians combined the visual appearance features reflected in the wire harness connector blind hole image data with destructive physical slicing verification (i.e., axially dissecting the wire harness connector to observe its internal structure under a microscope) and non-destructive airtightness testing results (i.e., applying a set pressure of gas to the connector and detecting the pressure drop rate to determine the actual sealing performance) for comprehensive evaluation. These two physical testing methods overcome the perspective limitations of single two-dimensional optical image judgment, enabling technicians to obtain the actual mechanical assembly state of the sealing bolt inside the blind hole and its final sealing function performance. This allows each sample to be labeled with an absolute physical truth value, generating a true label vector corresponding to the four business states mentioned above. (Uses one-hot encoding format).

[0046] Obtain the above real label vector Subsequently, it will serve as an objective basis for the network to evaluate the accuracy of its own prediction results. The network training process uses the cross-entropy loss function to calculate the prediction error, and its calculation formula is as follows: ; In the formula, This represents the cross-entropy loss value in the current iteration round; This indicates the total number of set sealing test status categories, which is 4 in this embodiment; Represents the th element in the true label vector. The true value of each category is labeled, taking the value of 0 or 1; This indicates the corresponding output of the defect classification module 400. The predicted probability values ​​for each category; This represents the natural logarithm operation. During training, the backpropagation algorithm of the detection system is used to calculate the gradient of the cross-entropy loss value with respect to the weights of all network layers. It is precisely because of the combination of physical slicing and airtightness testing in the early stages that high-fidelity real label vectors were generated. Only this cross-entropy loss value can truly reflect the model's current recognition defects, thus correctly guiding the descent direction of the weight gradient during backpropagation and preventing the model from learning incorrect visual features. Subsequently, an adaptive moment estimation optimizer is used to iteratively update the network parameters with a set initial learning rate. As the number of iteration batches increases, when the cross-entropy loss value on the validation set converges to a stable interval and no longer decreases, the training process ends and the final network weight model is saved, thereby ensuring accurate classification in the field testing phase.

[0047] Specific application examples: Take the waterproof wiring harness connector sealing performance testing station on an automotive manufacturing production line as an example. The wiring harness connector uses corrugated rubber sealing plugs internally, and the pre-calibrated physical depth scanning range of the blind holes in the wiring harness connector... Set to 15 mm.

[0048] The data acquisition module 100 drives the focusing plane of the programmable zoom lens to jump to a preset initial scanning position 2 mm above the upper edge of the blind aperture. The preset focal length step is set. It is 0.2 mm, according to the formula: Determine the total number of image frames acquired. The process involves 76 frames per second. The industrial camera triggers a global shutter at each set focal length plane to perform exposure imaging, acquiring a two-dimensional grayscale image at the corresponding depth of focus. The data acquisition module 100 sequentially overlays the 76 frames of 1024×1024 pixel two-dimensional grayscale images along spatial coordinates and focal length steps to construct a three-dimensional original focal stack image sequence.

[0049] The deep reconstruction module 200 extracts the 38th frame image from the original focal stack image sequence as the reference frame image. It then performs a two-dimensional discrete Fourier transform on the reference frame image and each frame to be registered in the original focal stack image sequence, calculates the cross-power spectrum, and extracts the extreme value coordinates as the translation compensation vector. In this embodiment, the maximum translation compensation vector extracted is 1.5 pixels horizontally and 0.8 pixels vertically. The deep reconstruction module 200 uses a bilinear interpolation algorithm to perform spatial interpolation translation on the original focal stack image sequence, outputting a spatially aligned focal stack sequence.

[0050] The depth reconstruction module 200 extracts gradient images using a 3×3 matrix Laplacian convolution kernel on the aligned focus stack sequence and calculates the local window variance using a 7×7 pixel local neighborhood window. Taking the center pixel at coordinates (512, 512) as an example, the 76 local window variance values ​​along the optical axis constitute a one-dimensional discrete curve. The depth reconstruction module 200 performs one-dimensional Gaussian fitting on this curve; the fitting principle is detailed in the appendix. Figure 3 As shown. In the appendix Figure 3 In the figure, the horizontal axis represents the focal length step number, and the vertical axis represents the local window variance. Figure 3 The discrete scatter points represent the actual local window variance data points extracted from each focal plane. The smooth solid line passing through the peak region of the scatter points is the generated Gaussian fitting continuous curve; the x-coordinate corresponding to the highest point of this Gaussian fitting continuous curve is the coordinate of the axis of symmetry. The width of the Gaussian fitted continuous curve represents the standard deviation. In this embodiment, the coordinates of the axis of symmetry of the fitted Gaussian curve are calculated. The standard deviation of the fitted Gaussian curve is 42.3. The value is 1.8. The deep reconstruction module 200 utilizes the symmetry axis coordinates. Multiply by focal length step (i.e., calculate 42.3 × 0.2 mm) Update the topological depth map, recording the sub-pixel physical depth corresponding to this pixel as 8.46 mm; Simultaneously, by performing linear interpolation along the optical axis on the corresponding pixels of the 42nd and 43rd frames, extract the grayscale value corresponding to the symmetry axis of the Gaussian curve, and perform pixel stitching in the spatial dimension to generate a fully focused texture map. The depth reconstruction module 200 extracts the standard deviation of the fitted Gaussian curve as the defocus gradient parameter, and stores 1.8 in the corresponding coordinate position of the defocus gradient parameter in the form of a two-dimensional matrix.

[0051] The strain calculation module 300 divides the generated 1024×1024 pixel fully focused texture map into 32×32 grid blocks in a spatial dimension, with each grid block having a size of 32×32 pixels. For the 150th grid block that has undergone physical compression deformation, the strain calculation module 300 performs a two-dimensional fast Fourier transform on it, shifts the zero-frequency component to the center of the spectrum, calculates the spectral energy outside the cutoff frequency, and generates high-frequency energy characteristics. The high-frequency energy characteristics calculated here are due to the abnormal compression of the rubber sealing plug by the inner wall of the blind hole, which causes the surface texture to become denser. The value is too high. The strain calculation module 300 retrieves the defocus gradient parameters in two-dimensional matrix form corresponding to grid block 150, calculates the average parameter value within the grid, and records it as the local defocus amount. Utilizing local defocus High-frequency energy characteristics The normalized spatial frequency for eliminating optical interference is obtained by performing a division correction calculation. The strain calculation module 300 retrieves the reference frequency distribution matrix pre-stored in the detection system. Calculate the relative rate of change between the two to generate a relative tensor value characterizing the degree of physical deformation. In this embodiment, the relative tensor value of the mesh... A value of 0.65, a positive value, reflects the high-frequency texture caused by the compression and contraction of the rubber sealing bolt at this location. The relative tensor values ​​of each grid block... The microscopic strain field tensor map is generated by arranging the coordinates of the two-dimensional grid.

[0052] The defect classification module 400 adjusts the topological depth map and fully focused texture map output from the depth reconstruction module 200 to 256×256 pixels using bilinear interpolation and directly concatenates them along the channel dimension to generate an initial bimodal tensor. This tensor is then input into a shallow feature extractor of a cross-modal feature fusion convolutional neural network, outputting a geometric feature map with a spatial resolution downsampled to 128×128 and a channel dimension expanded to 64. Simultaneously, the microscopic strain field tensor output from the strain calculation module 300 is also adjusted to a spatial resolution of 128×128 pixels using bilinear interpolation upsampling and input into the physical prior mapping network branch. After processing by a 1×1 convolutional layer and a sigmoid nonlinear activation function, an attention weight matrix with a dimension of 128×128×1 is generated. In the attention weight matrix, the aforementioned relative tensor values... The region with a value of 0.65 is converted into the corresponding high-weight pixel coordinates, with a weight value of 0.89.

[0053] The defect classification module 400 performs element-wise dot product operations on the geometric feature map and the attention weight matrix in the spatial dimension to generate a fused feature map. The fused feature map is input into the deep classification head network for feature dimensionality reduction and abstraction. After the fully connected layer calculates the logical values ​​for each classification, it is transformed into a classification probability distribution vector via the Softmax function. In this embodiment, the output classification probability distribution vector is [0.02, 0.03, 0.91, 0.04]. The defect classification module 400 uses the Argmax function to extract the index corresponding to the highest probability value of 0.91, maps it, and outputs the sealing detection result as "micro-flanged state," determining that the edge of the rubber sealing plug inside the wire harness connector is curled and stuck on the inner wall of the blind hole.

[0054] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A harness connector sealing property detection system based on machine vision and deep learning, characterized by, include: The data acquisition module is used to acquire the original focal stack image sequence of the wire harness connector through zoom scanning; The depth reconstruction module is used to perform phase registration on the original focal stack image sequence to obtain an aligned focal stack sequence, calculate the local window variance of the aligned focal stack sequence and perform Gaussian fitting to generate a topological depth map, defocus gradient parameters and full-focus texture map. The strain calculation module is used to calculate the high-frequency energy characteristics of the full-focus texture map, use the defocus gradient parameter to correct the high-frequency energy characteristics to obtain the normalized spatial frequency, and generate a micro-strain field tensor map based on the normalized spatial frequency. The defect classification module is used to extract the geometric feature maps of the topological depth map and the full-focus texture map, convert the micro-strain field tensor map into an attention weight matrix and perform weighted fusion with the geometric feature map to generate a fused feature map, and map the fused feature map to output the sealing detection result.

2. The machine vision and deep learning based wire harness connector sealing detection system of claim 1, wherein, The data acquisition module is equipped with an orthogonal polarization coaxial light source, a programmable zoom lens, and an industrial camera. The industrial camera drives the programmable zoom lens to continuously adjust the focal length according to a preset focal length step, and triggers a global shutter on each set focal length plane to perform exposure imaging and acquire two-dimensional grayscale images. The data acquisition module aligns and superimposes the continuously captured two-dimensional grayscale images along spatial coordinates with the preset focal length step to construct the original focal stack image sequence.

3. The machine vision and deep learning based wire harness connector sealing detection system of claim 1, wherein, The deep reconstruction module extracts the intermediate layer image from the original focal stack image sequence as the reference frame image. It performs a two-dimensional discrete Fourier transform on the reference frame image and each frame to be registered in the original focal stack image sequence to obtain a reference frequency domain complex matrix and a to-be-registered frequency domain complex matrix, respectively. It calculates the cross-power spectrum of the reference frequency domain complex matrix and the to-be-registered frequency domain complex matrix and performs an inverse Fourier transform to obtain the impulse response function in the spatial domain. It extracts the two-dimensional coordinates of the impulse response function at its extreme value as a translation compensation vector. Using a bilinear interpolation algorithm, it performs spatial interpolation translation on each frame to be registered based on the translation compensation vector, outputting the aligned focal stack sequence.

4. The machine vision and deep learning based wire harness connector sealing detection system of claim 1, wherein, The depth reconstruction module uses a Laplacian convolution kernel to perform convolution operations on each discrete focal length layer in the aligned focal stack sequence to obtain a gradient image. A local neighborhood window is set on the gradient image, and the variance of the gradient values ​​within the local neighborhood window is calculated as the local window variance. The focal length index corresponding to the maximum value of the local window variance is extracted along the optical axis. The focal length index is mapped to physical space depth by combining the preset focal length stride to generate the topological depth map.

5. The wire harness connector sealing performance testing system based on machine vision and deep learning according to claim 4, characterized in that, The depth reconstruction module performs one-dimensional Gaussian fitting on the curve formed by the local window variance along the optical axis to transform it into a continuous Gaussian curve function. It extracts the symmetry axis coordinates of the continuous Gaussian curve function, multiplies them by the preset focal length step size, and updates the topological depth map. It extracts the standard deviation of the continuous Gaussian curve function as the defocus gradient parameter. At the same time, it extracts the gray value corresponding to the symmetry axis coordinate by linear interpolating the corresponding pixels of adjacent discrete layer images along the optical axis and performs pixel stitching in the spatial dimension to generate the fully focused texture map.

6. The wire harness connector sealing performance testing system based on machine vision and deep learning according to claim 1, characterized in that, The strain calculation module divides the fully focused texture map into multiple non-overlapping grid blocks in a spatial dimension, performs a two-dimensional fast Fourier transform on each grid block, shifts the zero-frequency component to the center of the spectrum to obtain the frequency domain distribution matrix, calculates the spectral energy outside the set cutoff frequency radius, and generates the high-frequency energy features.

7. The machine vision and deep learning based wire harness connector sealing detection system of claim 6, wherein, The strain calculation module retrieves the defocus gradient parameters in two-dimensional matrix form corresponding to the grid block to calculate the local defocus amount. It then uses the local defocus amount to perform division compensation on the high-frequency energy characteristics to obtain the normalized spatial frequency. The module retrieves the pre-stored reference frequency distribution matrix to calculate the relative rate of change between the normalized spatial frequency and the reference frequency distribution matrix, generating a relative tensor value. Finally, the module arranges the relative tensor values ​​of each grid block into a two-dimensional matrix according to the two-dimensional grid coordinate sequence to generate the microscopic strain field tensor map.

8. The machine vision and deep learning based wire harness connector sealing detection system of claim 1, wherein, The defect classification module performs size normalization preprocessing on the topological depth map and the fully focused texture map, and directly concatenates them in the channel dimension to generate an initial bimodal tensor. The initial bimodal tensor is input into the shallow feature extractor of the cross-modal feature fusion convolutional neural network, and the local texture of the initial bimodal tensor is extracted through local receptive field sliding window operation to output the geometric feature map.

9. The machine vision and deep learning based wire harness connector sealing detection system of claim 8, wherein, The defect classification module adjusts the spatial resolution of the micro-strain field tensor map using bilinear interpolation upsampling. The micro-strain field tensor map with adjusted spatial resolution is input into the physical prior mapping network branch. The attention weight matrix is ​​generated through cross-channel integration and nonlinear activation function compression transformation. At the feature level, the geometric feature map and the attention weight matrix are multiplied element-wise in the spatial dimension to generate the fused feature map.

10. The machine vision and deep learning based wire harness connector sealing detection system of claim 9, wherein, The defect classification module inputs the fused feature map into a deep classification head network for feature dimensionality reduction and abstraction. The map is then compressed into a one-dimensional dense vector through a global average pooling layer and fed into a fully connected layer to calculate the logical values ​​of each category. The map is then converted into a classification probability distribution vector through an activation function. Finally, the index mapping corresponding to the highest probability value in the classification probability distribution vector is extracted and the sealing detection result is output.