Visual detection optimization control method, device and equipment based on digital twinning and storage medium

By employing frequency domain feature alignment and multi-scale fusion, the differences between virtual and real samples in local texture structure and channel distribution were resolved, improving the accuracy of visual inspection in identifying minute defects and the model's generalization ability, thus achieving higher detection accuracy and robustness.

CN121685996APending Publication Date: 2026-03-17SUZHOU HENGZHI INTELLIGENT TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511872278.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing visual inspection methods based on digital twins ignore local spatial structure constraints in the feature mapping between virtual and real samples, resulting in inaccurate texture alignment, which affects the accuracy of identifying minute defects and the generalization ability of the model. At the same time, speech spectrum feature alignment methods introduce channel correlation weakening and feature drift problems.

Method used

By using frequency domain feature alignment, multi-scale fusion, and cross-domain channel correction, the spectral features of simulated and real image samples are obtained, a frequency domain alignment model is established, structure-preserving alignment is performed, and features are extracted through multi-scale convolution and channel dependency mapping. The similarity between channels is quantified, and covariance regularization constraints are established to eliminate feature drift.

Benefits of technology

It significantly improves the recognition accuracy of visual inspection networks for subtle defects, edge textures, and abnormal surface gloss, enhances the model's perceptual sensitivity and robustness, and improves its adaptability under different production lines or data acquisition conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121685996A_ABST
    Figure CN121685996A_ABST
Patent Text Reader

Abstract

The invention discloses a visual detection optimization control method, device and equipment based on digital twinning and a storage medium, and relates to the technical field of visual detection, and the method comprises the steps: obtaining an analog image and a real image in a digital twinning environment, and carrying out the frequency domain feature transformation of the analog image and the real image, thereby obtaining an image spectrum feature; a frequency domain alignment model is established based on multi-band spectrum envelope guide residual mapping, and the structure of virtual and real image spectrum features is kept aligned; further extracting features through multi-scale convolution and channel dependence mapping to obtain virtual-real fusion features; quantifying channel similarity and establishing a covariance regularization constraint, performing channel correction on the cross-domain features, and eliminating feature drift to obtain second virtual-real fusion features; and finally, performing visual detection and micro defect identification based on the features. The problem that a virtual sample and a real sample are different in local texture structure and channel distribution is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of visual inspection, more particularly, the present application relates to a visual inspection optimization control method and device based on digital twinning, equipment and storage medium. BACKGROUND

[0002] The visual inspection optimization control method based on digital twinning has become an important means to improve the detection accuracy and automation level of industrial production lines. The prior art constructs a virtual production line model in a digital twinning environment, generates simulated image samples, and combines them with image data collected from the real production line to train and optimize the visual inspection model. This method can provide a large amount of controllable training data in the case of a lack of defect samples or high acquisition cost, and can assist in improving the performance and robustness of the detection model.

[0003] In a traditional digital twinning visual inspection system, in order to train a surface defect detection model, virtual image samples generated by simulation rendering are usually used. However, there are structural differences between the texture, lighting, and surface reflection of virtual samples and real captured images.

[0004] In existing visual inspection methods based on digital twinning, the feature mapping of virtual samples and real samples usually uses linear projection or shallow adversarial network structures, such as CycleGAN, Pix2Pix, etc., to achieve style transfer. However, when dealing with high-dimensional texture features, this kind of method usually only focuses on the overall similarity at the distribution level, ignoring the local spatial structure constraint, which can easily lead to structural distortion. The specific performance is that the texture details generated by the virtual sample domain are inconsistent with the local geometric morphology of the real image, which reduces the detection model's accuracy in identifying edge defects, micro-cracks, and glossy abnormalities, etc. This problem directly affects the micro-defect recognition accuracy and the model's generalization ability of the ultimate goal of visual inspection.

[0005] In order to alleviate the above-mentioned structural distortion problem, the feature mapping of virtual samples and real samples can be aligned more finely by analogy with the spectral feature alignment technology in the field of speech recognition. In speech recognition, by performing hierarchical matching and dynamic adjustment on spectral features, the local consistency of signals in time-frequency can be ensured, thereby improving the recognition accuracy. Similar ideas can be used in visual inspection to align the image spectral features in different frequency subbands to maintain the local texture consistency of virtual samples and real samples to some extent.

[0006] However, directly applying the speech spectrum feature alignment technology to image feature mapping also brings new cross-domain challenges. The spectrum of the speech signal mainly reflects the energy distribution of time-frequency, while the spectrum of the image has strong spatial local correlation. If a frequency domain feature mapping mechanism similar to speech is directly used, it may cause weakening of the correlation between channels in the convolutional features, causing feature drift and energy dispersion of local textures in the channel direction. This phenomenon will weaken the perception ability of the visual detection network to small textures and edge features, and even if the overall coherence of high-frequency textures is maintained, it may affect the accuracy of defect identification.

[0007] The existing visual detection optimization method based on digital twinning has obvious deficiencies: on the one hand, traditional linear mapping or shallow adversarial networks ignore local spatial structure constraints, resulting in inaccurate high-dimensional texture alignment between virtual samples and real samples; on the other hand, the spectrum feature alignment method analogous to speech recognition can maintain texture coherence to some extent, but introduces channel correlation weakening and feature drift problems in cross-domain feature mapping. Overall, existing technologies have not achieved fine consistency of virtual and real sample features in the frequency domain and spatial structure, limiting the performance of visual detection models in micro-defect identification.

[0008] In view of the above problems, the present application provides a solution. SUMMARY

[0009] In order to overcome the above-mentioned defects of the prior art, the embodiments of the present application provide a visual detection optimization control method, device, equipment and storage medium based on digital twinning, which performs frequency domain feature alignment, multi-scale fusion and cross-domain channel correction on simulated images and real images in a digital twinning environment to solve the problem of differences in local texture structure and channel distribution between virtual samples and real samples.

[0010] To achieve the above-mentioned purpose, the present application provides the following technical solutions: The visual detection optimization control method based on digital twinning comprises the following steps: acquiring simulated image samples and real image samples in a digital twinning environment, performing frequency domain feature transformation on the image samples to obtain image spectrum features; establishing a frequency domain alignment model of simulated image samples and real image samples based on multi-band spectrum envelope guided residual mapping to realize structure maintaining alignment of virtual and real image spectrum features; performing feature extraction on the aligned virtual and real image spectrum features based on multi-scale convolution and inserting channel dependent mapping to obtain virtual-real fusion features; quantifying the inter-channel similarity of the virtual-real fusion features, and establishing a covariance regularization constraint to correct the channel correlation of the cross-domain features, eliminate feature drift, and obtain second virtual-real fusion features; performing visual detection and defect identification based on the second virtual-real fusion features.

[0011] The device for a visual inspection optimization control method based on digital twins includes: a frequency domain transformation module, a structure alignment module, a virtual-real fusion module, a channel correlation correction module, and a visual inspection module. The frequency domain transformation module acquires simulated and real image samples in a digital twin environment, performs frequency domain feature transformation on the image samples, and obtains image spectral features. The structure alignment module establishes a frequency domain alignment model for simulated and real image samples based on multi-band spectral envelope-guided residual mapping according to the image spectral features, achieving structural alignment of virtual and real image spectral features. The virtual-real fusion module extracts features from the aligned virtual and real image spectral features based on multi-scale convolution and inserted channel dependency mapping, obtaining virtual-real fused features. The channel correlation correction module quantifies the inter-channel similarity of the virtual-real fused features, establishes covariance regularization constraints, corrects channel correlation for cross-domain features, eliminates feature drift, and obtains a second virtual-real fused feature. The visual inspection module performs visual inspection and defect identification based on the second virtual-real fused feature.

[0012] An electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform a visual inspection optimization control method based on digital twins.

[0013] A computer-readable storage medium storing a computer program that, when executed by a processor, implements a visual inspection optimization control method based on digital twins.

[0014] The technical effects and advantages of the present invention regarding the visual inspection optimization control method, device, equipment, and storage medium based on digital twins are as follows: 1. This invention acquires simulated and real image samples in a digital twin environment, and performs frequency domain feature transformation and multi-band spectral envelope-guided residual mapping on the image samples, achieving structural alignment between virtual and real samples in the frequency domain. Combining multi-scale convolution and channel-dependent mapping feature extraction methods, the generated virtual-real fusion features can simultaneously retain low-frequency structural information, mid-frequency texture features, and high-frequency detail information at different frequency levels. This technical solution effectively alleviates the problem of traditional linear projection or shallow adversarial networks focusing only on overall distribution similarity while ignoring local structural constraints, thereby significantly improving the recognition accuracy and model generalization ability of visual detection networks for targets such as subtle defects, edge textures, and abnormal surface gloss.

[0015] 2. This invention quantifies the inter-channel similarity of virtual-real fusion features and establishes covariance regularization constraints to correct channel correlation in cross-domain features, effectively eliminating inter-channel feature drift and obtaining a second virtual-real fusion feature with cross-domain consistency. When performing visual inspection and defect recognition based on this fusion feature, it ensures that the statistical structure between convolutional channels is consistent with the real sample, thereby maintaining the spatial and channel correlation of local textures and preventing feature energy dispersion. This technique not only enhances the detection model's sensitivity to minute defects, microcracks, and complex texture environments but also improves the adaptability and robustness of the entire system under different production lines or acquisition conditions. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating the visual inspection optimization control method based on digital twins according to the present invention.

[0017] Figure 2 This is a schematic diagram of the structure of the vision inspection optimization control device based on digital twin of the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0019] Example 1, Figure 1 The present invention provides a visual detection optimization control method based on digital twins, comprising the following steps: S1, acquire simulated image samples and real image samples in the digital twin environment, perform frequency domain feature transformation on the image samples to obtain image spectral features; In this embodiment, the step of acquiring simulated image samples and real image samples in a digital twin environment, and performing frequency domain feature transformation on the image samples to obtain image spectral features, specifically involves: Acquire real image samples of the inspection area of ​​the target production line, and establish a three-dimensional geometric constraint model based on the shooting parameters of the camera at the inspection station; Based on the three-dimensional geometric constraint model, simulated image samples consistent with the real perspective are generated in the digital twin simulation system to simulate typical defect morphology and form a simulated image sample set. The simulated image sample set and the real image sample set are subjected to size registration and illumination normalization processing to ensure that the two types of samples are consistent in spatial scale and brightness distribution. The normalized simulated image samples and real image samples are respectively input into the two-dimensional fast Fourier transform module to decompose and obtain the corresponding amplitude spectrum components and phase spectrum components. The image energy density curve for each frequency band is calculated based on the amplitude spectrum component, and the local texture structure distribution is extracted based on the phase spectrum component to construct a spectral feature matrix containing energy density and structural information.

[0020] It should be noted that the following are feasible examples of frequency band allocation and energy density calculation: ; In the formula, Let k be the energy density of the k-th frequency band. This is the set of frequency domain indices for the k-th frequency band. For the corresponding amplitude spectrum matrix elements, The number of elements within the frequency band.

[0021] It should be noted that the following is a computational example of feasible local texture structure extraction (phase distribution): Extract local texture gradients from the phase matrix within each frequency band: ; ; ; In the formula, Let be the local gradient of the frequency domain phase in the x-direction. Let be the local gradient of the frequency domain phase in the y-direction. For local texture intensity, The phase matrix for each frequency band.

[0022] It should be noted that the following is a feasible example of calculating the spectral characteristic matrix: By combining the energy density of each frequency band and the local texture intensity, the spectral feature matrix is ​​obtained: ; In the formula, The average local texture intensity of the k-th frequency band. It is a spectral feature matrix that includes energy and local structure.

[0023] The digital twin simulation system in this embodiment refers to establishing a three-dimensional virtual environment in a computer that is highly consistent with the actual production line inspection area. This virtual environment includes not only the spatial location of the inspection station, lighting conditions, and camera viewpoint, but also simulates the material reflection characteristics and the realistic distribution of defect morphologies of the target product. This system can generate virtual image samples that are consistent with the actual shooting perspective, making subsequent simulated samples comparable to real samples in terms of geometric structure and visual characteristics, thus providing fundamental data support for frequency domain feature alignment.

[0024] In this embodiment, the three-dimensional geometric constraint model refers to a spatial mapping model constructed by analyzing the shooting parameters (including focal length, pixel resolution, field of view, installation height, etc.) of the camera at the production line inspection station. This model can map the two-dimensional pixel coordinates of the actual captured image to three-dimensional space, while ensuring the consistency of the simulated image with the real image in terms of perspective, scale, and position in three-dimensional space. This ensures that defects in the simulated image exhibit the same spatial distribution characteristics as real defects in the virtual space.

[0025] In this embodiment, the amplitude spectrum component refers to the energy information of each frequency point obtained after performing a two-dimensional fast Fourier transform on the image, used to describe the overall brightness and texture energy distribution of the image. The phase spectrum component records the spatial arrangement and local shape features of the texture structure in the image. By simultaneously utilizing amplitude and phase information, not only the overall energy distribution of the image can be reflected, but also the spatial structure of local microscopic defects can be captured, achieving a comprehensive characterization of the features of both virtual and real images.

[0026] In this embodiment, the image energy density curve refers to the energy distribution curve obtained by statistically analyzing the amplitude spectrum in different frequency bands. It is used to measure the degree of energy concentration in the low-frequency, mid-frequency, and high-frequency ranges of the image. Local texture structure distribution refers to the spatial distribution patterns of local features such as texture, edges, microcracks, and unevenness in the image, extracted by analyzing the phase spectrum components. This information collectively forms the spectral feature matrix, ensuring that each image sample contains both energy and structural information, providing a complete data description for subsequent frequency domain alignment and feature fusion.

[0027] It should be noted that when generating simulated image samples, it is not enough to only ensure consistent viewpoint and simulated defect morphology. The simulated images must also undergo size registration and illumination normalization to ensure that their spatial scale, pixel distribution, and brightness distribution are consistent with the real images. This step ensures that the amplitude and phase spectra of the two types of images are comparable in subsequent frequency domain transformations, thereby effectively avoiding spectral shifts or structural distortions caused by differences in scale and brightness.

[0028] It should be noted that the spectral feature matrix is ​​not merely a simple arrangement of amplitude or phase spectra, but rather a unified mapping and combination of energy density curves and local texture structure distributions across different frequency bands, forming a comprehensive matrix containing both energy and structural information. This matrix provides clear frequency localization and local texture information during subsequent multi-band spectral envelope guidance and structure preservation processing, thereby ensuring point-by-point alignment and structural consistency of virtual and real spectral features.

[0029] S2, based on the image spectral features, establish a frequency domain alignment model for simulated image samples and real image samples using multi-band spectral envelope-guided residual mapping to achieve structural alignment of virtual and real image spectral features; In this embodiment, the step of establishing a frequency domain alignment model between simulated image samples and real image samples based on multi-band spectral envelope-guided residual mapping according to image spectral features specifically involves: The spectral feature matrices of simulated image samples and real image samples are divided into several sub-bands of low frequency, mid frequency and high frequency according to the frequency range, and the amplitude spectrum component and phase spectrum component in each sub-band are extracted respectively. The energy envelope function is calculated based on the amplitude spectrum components of each sub-band, and the envelope function is smoothed by interpolation to obtain the multi-band spectrum energy envelope curve. Within each sub-band, the energy difference vector between simulated image samples and real image samples is calculated based on the multi-band spectrum energy envelope curve, and a frequency band weight matrix is ​​constructed. The frequency band weight matrix is ​​used to perform band-weighted adjustment on the spectral feature matrix of the simulated image sample to correct the uneven energy distribution in different frequency bands and obtain the simulated spectral features after envelope guidance. The simulated spectral features guided by the envelope are combined with the spectral features of real image samples and used as the input dataset for the frequency domain alignment model.

[0030] It should be noted that the following are feasible calculation examples for frequency band division and sub-band extraction: The spectral feature matrix of the simulated image samples and the spectral feature matrix of the real image samples are divided into several sub-bands according to frequency range, such as low frequency, mid frequency, and high frequency. Each sub-band is a set of matrix indices. ; In the formula, This is the set of frequency indices for the k-th frequency band. , The spectral feature matrix of the simulated and real image samples; , Frequency domain index coordinates; , Let be the lowest and highest frequencies of the k-th frequency band.

[0031] It should be noted that the following is a feasible calculation example for calculating the frequency band amplitude energy envelope: For each sub-band, calculate the energy envelope function of the amplitude spectral component: ; By performing smooth interpolation on the envelope function, the multi-band spectral energy envelope curve is obtained: ; In the formula, To simulate the amplitude spectrum of image samples, The squared energy of the amplitude of the k-th subband; The smoothed energy envelope function; It is a smooth interpolation function.

[0032] It should be noted that the following is a feasible example of energy difference vector calculation: ; In the formula, Let be the energy difference matrix of the k-th subband.

[0033] It should be noted that the following is a calculation example of a feasible frequency band weight matrix construction: ; In the formula, Let be the weight matrix of the k-th sub-band. The absolute value of the energy difference. This represents the maximum energy difference within the subband.

[0034] In this embodiment, the spectral feature matrix sub-banding refers to dividing the entire spectrum into low-frequency, mid-frequency, and high-frequency sub-bands based on the frequency range of the image spectral feature matrix obtained after frequency domain transformation. The low-frequency sub-band mainly corresponds to the overall brightness distribution and large-scale structural information of the image, the mid-frequency sub-band corresponds to medium-scale texture and edge structures, and the high-frequency sub-band mainly reflects the image's minor defects, texture details, and local high-frequency information. Through this banding process, features in different frequency ranges can be independently analyzed and adjusted, achieving hierarchical energy correction and structure preservation.

[0035] In this embodiment, the multi-band spectral energy envelope curve refers to the curve formed by statistically analyzing the amplitude spectral components within each sub-band, calculating the energy density change at each frequency point, and then smoothing and interpolating these energy values ​​to create a continuous curve. This curve reflects the energy distribution trend of each sub-band and guides subsequent weighted adjustments to the spectral characteristics of the simulated image samples, making them closer to the real image samples in terms of overall energy distribution, thereby achieving the basic constraint of frequency domain alignment.

[0036] In this embodiment, the frequency band weighting matrix refers to the weighting coefficient matrix calculated based on the energy difference between the simulated image sample and the real image sample in each sub-band. This matrix is ​​used to perform band-weighted processing on the spectral features of the simulated image, focusing on enhancing the low-energy frequency bands and suppressing the high-energy frequency bands, thereby correcting the energy distribution of the simulated sample in the low, medium, and high-frequency sub-bands, so that its overall frequency domain structure is consistent with that of the real image sample.

[0037] In this embodiment, the simulated spectral features after envelope guidance refer to the simulated image sample spectral feature matrix after band-weighted adjustment using a frequency band weighting matrix. This matrix more closely approximates the real image sample at the amplitude and phase spectrum levels, while maintaining the spatial consistency of the local texture structure within each frequency sub-band, providing basic data for subsequent residual mapping and virtual-real feature fusion.

[0038] It should be noted that when performing subband partitioning and energy envelope calculation, it is essential to ensure that the frequency correspondence between the simulated image samples and the real image samples is strictly aligned. Otherwise, even with subsequent weighted adjustments, local high-frequency or low-frequency features may be distorted. Therefore, in this embodiment, subband partitioning and energy envelope calculation simultaneously consider frequency resolution and image sampling parameters to ensure complete comparability of virtual and real samples on the frequency axis.

[0039] It should be noted that the frequency band weighting matrix is ​​not only a tool for adjusting the energy of the amplitude spectrum, but also constrains the local structure of the phase spectrum to a certain extent. By weighting each sub-band, the texture direction and local edge structure of the simulated image in different frequency bands can be kept consistent with the real sample, avoiding structural misalignment or distortion of minor defects caused by energy imbalance, thus providing a reliable foundation for subsequent residual mapping and virtual-real fusion.

[0040] In this embodiment, the establishment of a frequency domain alignment model for simulated image samples and real image samples ensures the structural alignment of the virtual and real image spectral features, specifically as follows: Point-by-point difference operations are performed on the amplitude and phase components of the simulated spectral features after envelope guidance and the spectral features of the real image samples in the corresponding frequency bands to obtain the amplitude residual matrix and the phase residual matrix. The amplitude residual matrix and the phase residual matrix are combined into a frequency domain residual feature map in complex form, and maximum value normalization is performed in each frequency band to obtain a standardized frequency domain residual map. The gradient of the energy change rate at adjacent frequency points is calculated on the standardized frequency domain residual map, the local coherent gradient distribution is extracted, and the mean square error of the adjacent phase difference is calculated to obtain the global phase offset index. Based on the local coherent gradient distribution and the global phase shift index, a spectrum structure preservation mapping matrix is ​​constructed, and this matrix is ​​used to perform band-by-band correction on the simulated spectrum characteristics after envelope guidance. The simulated spectral features, after structure preservation correction, are weighted and fused with the spectral features of real image samples on the frequency axis to obtain a virtual-real frequency domain aligned feature set.

[0041] It should be noted that the following are feasible calculation examples for calculating the amplitude and phase difference point by point for the corresponding frequency band: ; ; In the formula, , Amplitude spectra for simulated and real samples; , Phase spectra of simulated and real samples; Let be the amplitude residual matrix of the k-th sub-band; Let be the phase residual matrix of the k-th subband.

[0042] It should be noted that the following are feasible examples of calculating local coherent gradient distribution and global phase shift: ; ; In the formula, For local magnitude gradient, This refers to the local phase mean square error. For the normalized frequency domain residual matrix, This is the standardized residual phase.

[0043] It should be noted that the following is a feasible example of calculating the normalized frequency domain residual matrix: ; In the formula, For the complex frequency domain residual of the k-th sub-band, This represents the maximum residual modulus within the subband.

[0044] In this embodiment, the amplitude residual matrix and phase residual matrix refer to the matrices obtained by comparing the simulated spectral features after envelope guidance with the spectral features of the real image samples point-by-point within the corresponding frequency band. The amplitude residual matrix reflects the difference in amplitude components at each frequency point, while the phase residual matrix reflects the shift in local phase information. Through this matrix representation, the differences in local structure and overall energy between virtual and real samples in the frequency domain can be refined, providing basic data for subsequent structure preservation correction.

[0045] In this embodiment, the frequency domain residual feature map refers to a two-dimensional frequency domain matrix synthesized from the amplitude residual matrix and the phase residual matrix in complex form, which can simultaneously contain energy difference and phase difference information. The standardized frequency domain residual map is the result after maximum value normalization within each sub-band, so that the differences between different frequency bands can be compared under the same dimension, ensuring that subsequent gradient calculation and structure mapping will not fail due to amplitude scale differences.

[0046] In this embodiment, the local coherent gradient distribution refers to the local gradient characteristics obtained by analyzing the energy change rate of adjacent frequency points on the standardized frequency domain residual map, reflecting the changing trend and texture consistency of the local spectral structure. The global phase offset index is an index obtained by statistically analyzing the mean square error of the phase difference between adjacent frequency points across the entire frequency range. It is used to evaluate the relative offset between virtual samples and real samples in the overall frequency domain structure, providing a constraint basis for the subsequent mapping matrix.

[0047] In this embodiment, the spectral structure-preserving mapping matrix refers to a mapping matrix constructed using local coherent gradient distribution and global phase offset indices, used to perform band-by-band correction on the simulated spectral features after envelope guidance. Through this mapping matrix, the local energy and phase distributions in the simulated sample spectrum can be adjusted to maintain structural consistency with the real sample, while preserving high-frequency minute texture information to the greatest extent possible, thereby achieving structural alignment between the virtual and real spectra.

[0048] In this embodiment, the virtual-real frequency domain aligned feature set refers to the feature set obtained by weighted fusion of the simulated spectral features corrected by the structure-preserving mapping matrix and the spectral features of the real image samples on the frequency axis. This feature set preserves the local texture details of the real samples while incorporating the diverse information of the simulated samples, providing a unified high-quality input for subsequent virtual-real fusion feature extraction and visual defect recognition.

[0049] It is important to note that when calculating amplitude and phase residuals, both local frequency differences and the overall frequency band structure must be considered. Focusing only on local differences may lead to over-adjustment of subtle high-frequency textures while distorting the low-frequency structure; conversely, considering only global phase shift may overlook local energy distribution anomalies. By jointly constructing a mapping matrix using local gradient distribution and global phase shift indices, both local and global factors can be taken into account, ensuring the integrity of the structural alignment.

[0050] It should be noted that during the weighted fusion of virtual and real frequency domain aligned feature sets, it is essential to ensure that the simulated samples and real samples are perfectly aligned on the frequency axis. Otherwise, the superposition of feature values ​​at different frequency points may lead to texture misalignment and energy drift. Therefore, in this embodiment, all spectral feature matrices undergo rigorous frequency registration and normalization to ensure that the fused features maintain a high-precision correspondence and structural consistency across the entire frequency range.

[0051] S3, based on multi-scale convolution and the insertion of channel dependency mapping, feature extraction is performed on the spectral features of the aligned virtual and real images to obtain virtual-real fusion features; In this embodiment, the step of extracting features from the aligned virtual and real image spectral features based on multi-scale convolution and the insertion of channel dependency mapping to obtain virtual-real fusion features is specifically as follows: The virtual and real frequency domain aligned feature sets are divided into several scale levels according to the frequency axis distribution, corresponding to the low-frequency structure layer, the mid-frequency texture layer and the high-frequency detail layer, respectively. Two-dimensional convolution operations are performed on the feature regions at each scale level, and spectral texture features under different spatial receptive fields are extracted according to the preset convolution kernel size to form a multi-scale feature mapping group. For each scale level of feature mapping group, calculate the global average response value and variance along the channel dimension, and construct the channel-dependent weight vector; Based on the channel dependency weight vector, the channel components of each scale feature mapping group are weighted and adjusted to enhance high-response channels and suppress low-correlation channels, resulting in a multi-scale feature set corrected by channel dependency mapping. The modified multi-scale feature set is fused and superimposed on the frequency axis and spatial axis, and the fusion result is standardized to obtain virtual-real fusion features that maintain frequency domain consistency and channel correlation.

[0052] In this embodiment, the virtual-real frequency domain aligned feature set refers to the set obtained by weighted fusion of simulated spectral features (corrected by the frequency domain structure-preserving mapping matrix) and real image sample spectral features on the frequency axis. This set contains diverse information from simulated samples and high-precision local texture information from real samples, providing a unified input for multi-scale convolution to extract features under different spatial receptive fields.

[0053] In this embodiment, the scale hierarchy refers to dividing the virtual and real frequency domain aligned feature set into a low-frequency structure layer, a mid-frequency texture layer, and a high-frequency detail layer according to the frequency axis, to reflect the differences in image texture and structure within different frequency ranges. The low-frequency structure layer mainly contains global contour information, the mid-frequency texture layer contains edges and texture patterns, and the high-frequency detail layer contains minute defects and fine texture information. Through multi-scale partitioning, feature extraction and processing at different frequency levels can be achieved.

[0054] In this embodiment, two-dimensional convolution refers to extracting local texture features using convolution kernels in the spectral feature regions at each scale level. The size of the convolution kernel can be preset to different spatial receptive fields to capture features ranging from local microstructures to larger texture patterns. The multi-scale feature mapping set refers to the set of output features after convolution at each scale level, including the mapping of different local patterns extracted by each convolution kernel, forming a comprehensive representation of the spectral features.

[0055] In this embodiment, the channel-dependent weight vector refers to the set of weights constructed by statistically analyzing the global average response value and variance of each scale feature mapping group along the channel dimension. This vector reflects the importance of each channel in representing virtual-real fusion features. By adjusting the weights, channels with high information density can be enhanced, while channels with low or irrelevant responses can be suppressed, thereby optimizing feature representation.

[0056] In this embodiment, the multi-scale feature set corrected by channel-dependent mapping refers to the set obtained after applying a channel-dependent weight vector to weight and adjust the channel components of the multi-scale feature mapping group. This set takes into account the important features of each frequency level, while eliminating information redundancy and weakening between channels, providing highly consistent and discriminative feature inputs for subsequent frequency axis and spatial axis fusion.

[0057] In this embodiment, the virtual-real fusion feature refers to the final feature set obtained by fusing and superimposing the multi-scale feature set after channel dependency mapping correction on the frequency and spatial axes and then standardizing it. This feature set maintains consistency in both the frequency and spatial domains, and the correlation between channels is optimized, effectively representing the joint features of virtual and real samples, providing a unified and high-quality input for subsequent cross-domain feature correction and defect detection.

[0058] It should be noted that when using multi-scale convolution to extract features, information imbalance may occur between different scale levels. That is, low-frequency features may be masked by high-frequency features, and high-frequency subtle textures may be weakened by low-frequency information. By introducing channel-dependent weight vectors and performing weighted adjustments, we can ensure that the importance of each channel is reasonably reflected while maintaining the integrity of features at each scale, thus achieving a balanced multi-scale representation of features.

[0059] It should be noted that after weighted adjustment of the multi-scale feature mapping group, directly using a single-scale feature for detection may result in insufficient information or local texture distortion in the virtual-real fusion features at certain frequency levels. By fusing and overlaying the features on the frequency and spatial axes and performing standardization, the local and global information of the features at each scale can be integrated, making the virtual-real fusion features more complete and consistent in cross-domain feature representation, and providing reliable input for subsequent covariance regularization and defect identification.

[0060] S4 quantifies the inter-channel similarity of virtual-real fusion features and establishes covariance regularization constraints to correct channel correlation for cross-domain features, eliminate feature drift, and obtain the second virtual-real fusion feature.

[0061] In this embodiment, the quantification of the inter-channel similarity of the virtual-real fusion features specifically refers to: The virtual-real fusion features are expanded along the channel dimension, and the feature response vector of each channel is extracted to form a channel feature set. Calculate the cosine similarity value for the feature response vectors of any two channels in the channel feature set, and construct a channel similarity matrix using the channel number as an index; The channel weight coefficient vector is calculated based on the average frequency band energy of each channel, and the channel similarity matrix is ​​weighted and corrected to obtain the weighted channel similarity matrix.

[0062] In this embodiment, the channel feature set refers to the set of feature response vectors corresponding to each channel after expanding the virtual-real fusion features along the channel dimension. Each vector reflects the comprehensive response information of that channel on the frequency axis and spatial axis, including low-frequency structure, high-frequency texture, and minor defect features. By forming the channel feature set, the correlation between different channels can be quantitatively analyzed, providing basic data for subsequent similarity calculation and channel correction.

[0063] In this embodiment, cosine similarity is an index used to measure the similarity in the directions of the feature response vectors of any two channels. Its value reflects the consistency of the two channels in the virtual-real fusion feature space. The channel similarity matrix is ​​a set of matrix-like structures that organizes the pairwise similarities between all channels using the channel number as an index. Each element in the matrix represents the magnitude of the response consistency between the corresponding two channels, which is used to guide channel correlation adjustment and feature drift elimination.

[0064] In this embodiment, the channel weight coefficient vector refers to the set of weights calculated based on the average energy of the corresponding frequency band of each channel. This vector reflects the importance of each channel in the overall feature representation. By weighting, the contribution of high-energy channels can be enhanced, while low-energy or redundant channels can be suppressed, thereby optimizing cross-domain feature representation and stability.

[0065] In this embodiment, the weighted channel similarity matrix is ​​a matrix obtained by weighting and correcting each element using the channel weight coefficient vector based on the original channel similarity matrix. This matrix not only reflects the consistency of responses between channels but also considers the differences in the contribution of each channel to the overall features, providing a precise basis for covariance regularization and channel correlation correction.

[0066] It should be noted that the virtual-real fusion features contain multi-scale and multi-channel information along the frequency and spatial axes. Directly processing the overall features can easily overlook the local consistency and cross-domain differences between channels. By expanding along the channel dimension to form a channel feature set and performing cosine similarity quantification, the similarity and deviation between channels can be accurately captured, providing data support for subsequent channel correction.

[0067] It should be noted that using only the original channel similarity matrix for correction may lead to low-energy or redundant channels having an unnecessary impact on the overall correction. By calculating the average frequency band energy of each channel and forming a channel weight coefficient vector, the channel similarity matrix can be weighted and corrected to ensure that high-response channels are fully preserved and optimized. This allows cross-domain feature correction to both eliminate drift and maintain the effective information integrity of virtual-real fusion features.

[0068] In this embodiment, the establishment of covariance regularization constraints, channel correlation correction of cross-domain features, and elimination of feature drift to obtain the second virtual-real fusion feature are specifically as follows: Based on the weighted channel similarity matrix, calculate the covariance matrix of the simulated channel set and the real channel set respectively; Calculate the difference matrix between the virtual and real channel covariance matrices, and then standardize the difference matrix according to the channel dimension to obtain the standardized covariance offset matrix; The standardized covariance offset matrix is ​​applied to the channel components of the virtual-real fusion feature, and linear remapping correction is performed on the feature response of each channel to correct the correlation deviation between channels. The corrected channel features are normalized channel by channel to ensure that the mean and variance of each channel are consistent with the true channel. The normalized correction features are recombined to form the second virtual-real fusion feature, and then fused and superimposed on the frequency axis and spatial axis to obtain the second virtual-real fusion feature with cross-domain channel consistency and drift elimination.

[0069] In this embodiment, the weighted channel similarity matrix refers to a matrix that is weighted and corrected based on the channel similarity of the virtual-real fusion features, combined with the frequency band energy weights of each channel. Each element in the matrix reflects the degree of consistency of the responses of the corresponding two channels in the cross-domain feature space. This matrix is ​​used to guide subsequent covariance calculation and channel correction to ensure that the correction process focuses on important channels and reduces interference from low-correlation channels.

[0070] In this embodiment, the covariance matrix is ​​obtained by calculating the statistical correlation between the feature responses of each channel in the channel set of the simulated image samples and the channel set of the real image samples, respectively. The matrix elements describe the joint variation trend and correlation strength of any two channels, reflecting the overall dependency structure between channels and providing basic information for eliminating cross-domain feature drift.

[0071] In this embodiment, the normalized covariance offset matrix refers to the matrix obtained by normalizing the difference between the covariance matrices of the simulated channel set and the real channel set by channel. The matrix elements represent the magnitude and direction of each channel's deviation from the true distribution in cross-domain features, and are used for subsequent channel-level linear remapping and correction of the virtual-real fusion features.

[0072] In this embodiment, the channel linear remapping correction refers to linearly adjusting the feature response of each channel according to the standardized covariance offset matrix, so that the correlation between channels is consistent with that of the real channels. The normalization process further ensures that the mean and variance of each channel after correction are aligned with the real channels, thereby forming a cross-domain consistent second virtual-real fusion feature and eliminating the drift between virtual and real features.

[0073] It should be noted that during the virtual-real fusion feature extraction process, there are statistical distribution differences in the channel responses of simulated image samples and real image samples, leading to cross-domain feature drift. Without channel correlation correction and covariance constraints, the accuracy of subsequent defect identification will be affected. By establishing covariance regularization constraints and performing linear remapping, inter-channel bias can be effectively eliminated, ensuring that the statistical properties of the virtual-real fusion features are consistent with those of the real channels.

[0074] It should be noted that directly using covariance differences for correction may lead to over-correction of high-energy channels and under-correction of low-energy channels. Standardization normalizes the difference matrix by channel, ensuring that each channel is adjusted appropriately according to its deviation during the correction process. This guarantees the balance of the correction process while maintaining cross-domain consistency and integrity of features, providing stable and reliable input features for final defect identification.

[0075] S5 performs visual inspection and defect recognition based on the second virtual-real fusion feature.

[0076] In this embodiment, the visual detection and defect identification based on the second virtual-real fusion feature specifically includes: The second virtual-real fusion feature is input into the existing convolutional neural network visual detection model to extract multi-level spatial and frequency information. The extracted features are used for target localization and classification, and the defect location coordinates and defect type probability distribution are output. Valid defects are filtered based on the defect coordinates and probability thresholds of the detection output, and a defect annotation map and defect list are generated. The defect annotation map is overlaid on the original image for visual verification and manual review of the inspection results; The defect list, along with corresponding production station information, acquisition time, and other metadata, is recorded in the detection database for subsequent feature mapping updates and model adaptive optimization in the digital twin environment.

[0077] Example 2, Figure 2 The present invention provides an apparatus for a visual inspection optimization control method based on digital twins, comprising: a frequency domain transformation module, a structure alignment module, a virtual-real fusion module, a channel correlation correction module, and a visual inspection module; the frequency domain transformation module is used to acquire simulated image samples and real image samples in a digital twin environment, and perform frequency domain feature transformation on the image samples to obtain image spectral features; the structure alignment module is used to establish a frequency domain alignment model of simulated image samples and real image samples based on multi-band spectral envelope guided residual mapping according to the image spectral features to achieve structural alignment of virtual and real image spectral features; the virtual-real fusion module is used to extract features from the aligned virtual and real image spectral features based on multi-scale convolution and inserting channel dependency mapping to obtain virtual-real fused features; the channel correlation correction module is used to quantify the inter-channel similarity of virtual-real fused features, establish covariance regularization constraints, perform channel correlation correction on cross-domain features, eliminate feature drift, and obtain a second virtual-real fused feature; the visual inspection module is used to perform visual inspection and defect identification based on the second virtual-real fused feature.

[0078] This embodiment also includes an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to execute a visual inspection optimization control method based on digital twins.

[0079] This embodiment also includes a computer-readable storage medium storing a computer program that, when executed by a processor, implements a visual detection optimization control method based on digital twins.

[0080] In the embodiments provided by this invention, it should be understood that the disclosed system can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.

[0081] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.

[0082] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the invention. No appended diagram markings in the claims should be construed as limiting the scope of the claims.

[0083] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0084] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0085] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above.

[0086] In the embodiments provided in this disclosure, it should be understood that the disclosed systems and methods can also be implemented in other ways. The system embodiments described above are merely illustrative; for example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0087] It should be noted that, in this disclosure, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element limited by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0088] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A visual inspection optimization control method based on digital twinning, characterized in that, The method comprises the following steps: Obtaining simulation image samples and real image samples in a digital twin environment, and performing frequency domain feature transformation on the image samples to obtain image spectrum features; According to the image spectrum features, a frequency domain alignment model of the simulation image samples and the real image samples is established based on multi-band spectrum envelope guided residual mapping to realize structure maintaining alignment of virtual and real image spectrum features; Based on multi-scale convolution and insertion of channel-dependent mapping, feature extraction is performed on the aligned virtual and real image spectrum features to obtain virtual-real fusion features; The inter-channel similarity of the virtual-real fusion features is quantified, and a covariance regularization constraint is established to correct the channel correlation of the cross-domain features, eliminate feature drift, and obtain second virtual-real fusion features; Based on the second virtual-real fusion features, visual inspection and defect recognition are performed.

2. The digital-twin-based visual inspection optimization control method of claim 1, wherein, The method comprises the following steps: Obtaining simulation image samples and real image samples in a digital twin environment, and performing frequency domain feature transformation on the image samples to obtain image spectrum features, specifically: Obtaining real image samples of a target production line detection area, and establishing a three-dimensional geometric constraint model based on the shooting parameters of the detection station camera; According to the three-dimensional geometric constraint model, simulation image samples consistent with the real visual angle are generated in the digital twin simulation system, typical defect morphologies are simulated, and a simulation image sample set is formed; The simulation image sample set and the real image sample set are subjected to size registration and illumination normalization processing to maintain consistency in spatial scale and brightness distribution between the two types of samples; The normalized simulation image samples and real image samples are respectively input into a two-dimensional fast Fourier transform module to obtain corresponding amplitude spectrum components and phase spectrum components; 3. The digital-twin-based visual inspection optimization control method of claim 2, wherein, Based on the amplitude spectrum components, the image energy density curve of each frequency band is calculated, and based on the phase spectrum components, the local texture structure distribution is extracted to construct a spectrum feature matrix containing energy density and structure information. The method comprises the following steps: The spectrum feature matrices of the simulation image samples and the real image samples are divided into low-frequency, medium-frequency and high-frequency sub-bands according to the frequency range, and the amplitude spectrum components and the phase spectrum components in each sub-band are extracted; Based on the amplitude spectrum components of each sub-band, an energy envelope function is calculated, and the envelope function is subjected to smoothing interpolation to obtain a multi-band spectrum energy envelope curve; In each sub-band, an energy difference vector between the simulation image samples and the real image samples is calculated according to the multi-band spectrum energy envelope curve, and a frequency band weight matrix is constructed; The spectrum feature matrix of the simulation image samples is subjected to sub-band weighting adjustment using the frequency band weight matrix to correct the uneven energy distribution of different frequency bands and obtain envelope-guided simulation spectrum features; 4. The digital-twin-based visual inspection optimization control method of claim 3, wherein, The envelope-guided simulation spectrum features and the spectrum features of the real image samples are combined as the input data set of the frequency domain alignment model. The method comprises the following steps: The amplitude component and the phase component of the envelope guided simulation spectrum feature and the spectrum feature of the real image sample in the corresponding frequency band are subjected to point-by-point difference operation to obtain an amplitude residual matrix and a phase residual matrix; The amplitude residual matrix and the phase residual matrix are combined into a complex frequency domain residual feature map, and a maximum value normalization processing is performed in each frequency band to obtain a standardized frequency domain residual map; Gradient calculation is performed on the energy change rate of adjacent frequency points on the standardized frequency domain residual map to extract a local coherent gradient distribution, and a mean square error of adjacent phase difference is calculated to obtain a global phase offset index; Based on the local coherent gradient distribution and the global phase offset index, a spectrum structure preservation mapping matrix is constructed, and the envelope guided simulation spectrum feature is corrected in each frequency band by using the matrix; The simulation spectrum feature corrected by the structure preservation is fused with the spectrum feature of the real image sample on the frequency axis to obtain a virtual-real frequency domain alignment feature set.

5. The digital-twin-based visual inspection optimization control method of claim 4, wherein, The virtual-real image spectrum features after alignment are extracted by using the multi-scale convolution and inserting channel-dependent mapping to obtain virtual-real fusion features, specifically as follows: The virtual-real frequency domain alignment feature set is divided into several scale levels according to the frequency axis distribution, which correspond to a low-frequency structure layer, a medium-frequency texture layer and a high-frequency detail layer respectively; Two-dimensional convolution operations are respectively performed on the feature regions of each scale level to extract spectrum texture features under different spatial receptive fields according to a preset convolution kernel size, and a multi-scale feature mapping group is formed; Global average response values and variances are calculated along the channel dimension for the feature mapping group of each scale level to construct a channel-dependent weight vector; Based on the channel-dependent weight vector, the channel components of each scale feature mapping group are weighted and adjusted to enhance high-response channels and suppress low-relevance channels, and a multi-scale feature set corrected by channel-dependent mapping is obtained; The corrected multi-scale feature set is fused and superimposed on the frequency axis and the spatial axis, and the fusion result is standardized to obtain virtual-real fusion features that maintain frequency domain consistency and channel relevance.

6. The digital-twin-based visual inspection optimization control method of claim 5, wherein, The inter-channel similarity of the quantized virtual-real fusion features is specifically as follows: The virtual-real fusion features are unfolded along the channel dimension to extract feature response vectors of each channel, and a channel feature set is formed; Cosine similarity values are calculated for the feature response vectors of any two channels in the channel feature set, and a channel similarity matrix is constructed with the channel number as the index; A channel weight coefficient vector is calculated based on the average energy of each channel, and the channel similarity matrix is weighted and corrected to obtain a weighted channel similarity matrix.

7. The digital-twin-based visual inspection optimization control method of claim 6, wherein, The covariance regularization constraint is established to correct the channel correlation of the cross-domain features and eliminate feature drift to obtain second virtual-real fusion features, specifically as follows: Covariance matrices of the simulation channel set and the real channel set are respectively calculated according to the weighted channel similarity matrix; A difference matrix of the virtual-real channel covariance matrices is calculated, and the difference matrix is standardized along the channel dimension to obtain a standardized covariance offset matrix; The standardized covariance offset matrix is applied to the channel components of the virtual-real fusion features to linearly remap and correct the feature response of each channel and correct the correlation deviation between channels. The corrected channel features are normalized channel by channel, so that the mean and variance of each channel are consistent with the real channel; The normalized corrected features are recombined to form second virtual-real fusion features, and are fused and superimposed on the frequency axis and the spatial axis to obtain second virtual-real fusion features that are consistent across domains and eliminate drift.

8. An apparatus using the digital-twin-based visual inspection optimization control method according to any one of claims 1-7, characterized in that, Comprise: a frequency domain conversion module, a structure alignment module, a virtual-real fusion module, a channel correlation correction module, and a visual detection module; The frequency domain conversion module is configured to obtain simulation image samples and real image samples in a digital twin environment, and perform frequency domain feature transformation on the image samples to obtain image spectrum features; The structure alignment module is configured to establish a frequency domain alignment model of the simulation image samples and the real image samples based on multi-band spectral envelope guided residual mapping according to the image spectrum features, to realize structure maintaining alignment of virtual and real image spectrum features; The virtual-real fusion module is configured to perform feature extraction on the aligned virtual and real image spectrum features based on multi-scale convolution and insertion of channel dependent mapping to obtain virtual-real fusion features; The channel correlation correction module is configured to quantify the inter-channel similarity of the virtual-real fusion features, and establish a covariance regularization constraint to correct the channel correlation of the cross-domain features and eliminate feature drift to obtain second virtual-real fusion features; The visual detection module is configured to perform visual detection and defect recognition based on the second virtual-real fusion features.

9. An electronic device, comprising: The electronic device comprises: at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the digital twin-based visual detection optimization control method of any one of claims 1 to 7.

10. A computer readable storage medium storing a computer program, characterized in that, The computer program is executed by the processor to implement the digital twin-based visual detection optimization control method of any one of claims 1 to 7.

Citation Information

Cited By

  • Architectural coating detection method based on infrared polarization imaging

    CN121861042A

  • A method for detecting a building coating based on infrared polarization imaging

    CN121861042B