Visual detection method and device, computer equipment and storage medium
By combining illumination compensation, noise reduction, and contrast enhancement with multi-layer feature extraction from convolutional neural networks, the accuracy problem of target recognition in dynamic and complex environments in visual inspection is solved, achieving efficient and stable industrial visual inspection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 彩迅工业(中山)有限公司
- Filing Date
- 2025-12-19
- Publication Date
- 2026-05-08
AI Technical Summary
Existing visual inspection methods struggle to accurately identify targets in dynamic and complex processing environments, especially when faced with image variations under different workstations and lighting conditions, which challenges the accuracy and robustness of the inspection.
By acquiring real-time image data, performing illumination compensation and noise reduction, and combining contrast enhancement, a pre-trained convolutional neural network is used to extract multi-layer features, train a target detection model, and achieve accurate identification of targets in the image.
Maintaining high detection accuracy and stability in complex production environments enables accurate differentiation of multiple types of targets, meeting the needs of efficient, real-time, and highly reliable industrial vision inspection.
Smart Images

Figure CN121999271A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of visual processing technology, specifically relating to a visual inspection method, apparatus, computer equipment, and storage medium. Background Technology
[0002] Industrial visual inspection, as a crucial pillar of modern manufacturing, bears the key mission of improving product quality and production efficiency. In the context of intelligent manufacturing, the precise identification of target products through image analysis technology can not only reduce the cost of manual inspection but also significantly improve the stability and reliability of inspection. Research and application in this field are directly related to the level of intelligence in industrial production and the core competitiveness of enterprises.
[0003] However, many current visual inspection methods face significant limitations in practical applications. Traditional inspection methods often rely on fixed rules or preset templates, making it difficult to adapt to complex and ever-changing processing environments. This limitation manifests in insufficient ability to capture image features during dynamic production processes, especially when faced with image variations at different workstations and under different lighting conditions, where the accuracy and robustness of the inspection are often challenged.
[0004] Taking logo detection as an example, traditional methods require pre-setting the standard size, angle, and grayscale features of the logo. When the logo is rotated, scaled, or partially obscured due to positioning deviations on the production line, the similarity score of template matching will drop significantly, making it very easy to miss or misdetect the logo.
[0005] Therefore, how to accurately identify the target to be inspected from real-time captured images in a dynamic and complex processing environment has become a key problem that urgently needs to be solved in the field of industrial vision inspection. Summary of the Invention
[0006] The purpose of this application is to provide a visual inspection method, apparatus, computer device, and storage medium to solve the technical problem of accurately identifying and detecting targets from real-time captured images in a dynamic and complex processing environment.
[0007] To address the aforementioned technical problems, this application provides a visual detection method, employing the following technical solution: A visual inspection method, comprising: The first image is obtained by acquiring raw image data captured in real time from the processing environment. To address lighting variations and background interference in the processing environment, the first image is denoised to obtain the second image. The second image is contrast-enhanced to obtain the third image; Obtain the pixel distribution information of the third image, and use a pre-trained convolutional neural network based on the pixel distribution information to extract multi-layer features from the third image, thereby obtaining a set of feature vectors containing multiple visual morphological details. The feature vector set is used as the data training set, and the preset initial classification model is trained using the data training set to obtain the object detection model; Receive detection instructions, acquire the image to be detected, and perform target detection on the image to be detected based on the target detection model to obtain the target detection result.
[0008] To address the aforementioned technical problems, this application also provides a visual inspection device, which employs the following technical solution: A visual inspection device, comprising: The real-time imaging module is used to acquire raw image data captured in real time from the processing environment to obtain the first image; The noise reduction module is used to denoise the first image to obtain the second image, taking into account changes in lighting and background interference in the processing environment. A contrast enhancement module is used to enhance the contrast of the second image to obtain a third image; The feature extraction module is used to obtain the pixel distribution information of the third image. Based on the pixel distribution information, a pre-trained convolutional neural network is used to extract multi-layer features from the third image to obtain a set of feature vectors containing multiple visual morphological details. The classification training module is used to use the feature vector set as the data training set and to train the preset initial classification model using the data training set to obtain the object detection model; The target detection module is used to receive detection instructions, acquire the image to be detected, and perform target detection on the image to be detected based on the target detection model to obtain the target detection result.
[0009] To address the aforementioned technical problems, this application also provides a computer device that employs the following technical solution: A computer device includes a memory and a processor, the memory storing computer-readable instructions, the processor executing the computer-readable instructions to implement the steps of the visual inspection method as described in any of the preceding claims.
[0010] To address the aforementioned technical problems, this application also provides a computer-readable storage medium, employing the technical solution described below: A computer-readable storage medium storing computer-readable instructions that, when executed by a processor, implement the steps of the visual inspection method as described in any one of the preceding descriptions.
[0011] Compared with the prior art, the embodiments of this application have the following main advantages: This application discloses a visual inspection method, apparatus, computer equipment, and storage medium. It effectively improves image quality through preprocessing steps such as illumination compensation, noise reduction, and background suppression. Contrast enhancement further amplifies local grayscale differences in potential target areas, making previously difficult-to-identify subtle features such as dents, scratches, and cracks more prominent in the image. The introduction of a pre-trained convolutional neural network, combined with multi-scale feature extraction, residual enhancement, and feature fusion, enables the system to simultaneously capture visual features of targets of different sizes and depth levels. By using the extracted multi-level effective feature vectors to train a classification model, the system can gradually develop target recognition capabilities for specific processed workpieces, achieving accurate differentiation of multiple target types. In the real-world inspection stage, the model rapidly infers from the image to be inspected, maintaining high detection accuracy and stability in complex production environments, thus meeting the demands of efficient, real-time, and highly reliable industrial visual inspection. Attached Figure Description
[0012] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 An exemplary system architecture diagram is shown, in which this application can be applied; Figure 2 A flowchart of one embodiment of the visual inspection method according to this application is shown; Figure 3 It shows Figure 2 A flowchart of an embodiment of step S204; Figure 4 This invention provides a schematic diagram of the structure of a display device logo inspection station according to a specific embodiment of the present application. Figure 5 A schematic diagram of the structure of one embodiment of the visual inspection device according to this application is shown; Figure 6 It shows Figure 5 A schematic diagram of a embodiment of the feature extraction module 404; Figure 7 A schematic diagram of the structure of one embodiment of a computer device according to this application is shown. Detailed Implementation
[0014] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.
[0015] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0016] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0017] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables.
[0018] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.
[0019] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.
[0020] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.
[0021] It should be noted that the visual inspection method provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the visual inspection device is generally set in the server / terminal device.
[0022] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative; the system can have any number of terminal devices, networks, and servers depending on implementation needs.
[0023] Continue to refer to Figure 2 A flowchart of one embodiment of the visual inspection method according to this application is shown. The visual inspection method includes the following steps: S201, acquire raw image data captured in real time from the processing environment to obtain the first image; Specifically, industrial cameras, intelligent acquisition terminals, or image acquisition modules with high dynamic range imaging capabilities, deployed at different locations within the processing area, continuously capture images of the target workpiece, dynamically changing areas that may occur during processing, and the surrounding background. To ensure the effectiveness of the acquired images, the system can automatically adjust the acquisition frame rate and shutter speed according to the operating status of the processing equipment. For example, it can increase the sampling frequency during high-speed processing or rapid tool movement to avoid image blurring. Furthermore, imaging parameters, including exposure time, gain, white balance, and aperture size, can be controlled in real time during acquisition to keep the image's grayscale distribution within a manageable range. To adapt to different processing scenarios, such as highly reflective metal surfaces, rough black workpieces, or slippery surfaces with liquid coolant, polarizing filters, supplementary lighting modules, or adjustable light source layouts can be used to reduce light interference and improve imaging stability. The acquired images can be output as is or undergo preliminary processing via the camera's internal ISP (Image Signal Processing) module, such as color correction, noise reduction, and distortion correction, before being used as the final image.
[0024] S202, denoising the first image to address lighting changes and background interference in the processing environment, thereby obtaining the second image; Specifically, the process begins with a global brightness analysis of the image. By statistically analyzing the mean, variance, and peak grayscale distribution of image pixels in different regions, it's possible to determine if there are significant overexposure, underexposure, or uneven local lighting. Based on this, brightness reconstruction is performed on areas with strong light reflection and shadow to mitigate the impact of uneven lighting on feature recognition. Subsequently, to address common noise issues in the processing environment, such as particulate noise, texture interference from environmental dust, and random noise generated under high camera sensitivity, spatial domain filtering (e.g., median filtering, bilateral filtering), frequency domain filtering (e.g., bandpass filtering, wavelet denoising), or self-learning denoising networks are employed to eliminate invalid noise while preserving the image's edge structure. Furthermore, background texture stripping can be achieved using background modeling methods, such as Gaussian background models, background differencing, and structured texture suppression, making the workpiece area more prominent in the image. The final output image exhibits lower noise levels and weaker background interference.
[0025] S203, perform contrast enhancement on the second image to obtain a third image, wherein the visual morphology of potential targets in the image is highlighted by contrast enhancement; Specifically, the second image is first analyzed using a grayscale histogram to statistically analyze the distribution of pixels at each grayscale level, identifying areas of concentrated grayscale and regions with narrow dynamic range. For images with local contrast differences, such as fine scratches on a workpiece surface with similar grayscale to the surrounding area, adaptive enhancement is applied to the local areas to amplify the grayscale difference between the potential target area and the background. During enhancement, to avoid over-enhancement leading to noise amplification, upper and lower thresholds for grayscale transformation are set, and pixels exceeding the threshold range are truncated. Simultaneously, an edge protection mechanism is used to ensure that the edge structure of the target is not damaged while enhancing contrast. After contrast enhancement, the contour features, surface texture details, and visual morphology of potential target areas in the third image are significantly highlighted.
[0026] S204, obtain the pixel distribution information of the third image, and extract multi-layer features from the third image using a pre-trained convolutional neural network based on the pixel distribution information to obtain a set of feature vectors containing multiple visual morphological details; Specifically, the third image is first analyzed and statistically analyzed for pixel density, including pixel brightness distribution, texture orientation distribution, and gradient variation range, to generate pixel distribution description information that can guide feature extraction by a convolutional neural network. Then, the third image is input into a pre-trained convolutional neural network model, such as the YOLO series, ResNet, Darknet, or other CNN models with multi-scale feature response structures. The pre-trained model can utilize its learned general visual feature representation capabilities to perform multi-layer convolutional operations on the input image, including shallow edge feature extraction, mid-layer texture structure recognition, and deep semantic information encoding. During feature extraction, the model's internal feature pyramid structure, residual connections, attention mechanisms, or feature fusion strategies can further enhance the response capability for recognizing subtle targets, enabling the network to simultaneously recognize details at different scales and in different directions. In the outputs of all convolutional layers, feature layers, and fusion layers, the system reduces the dimensionality of relevant feature maps, encodes them, or serializes them, ultimately generating a set of feature vectors containing various detailed information such as target contours, texture features, shape features, and edge structures.
[0027] In one specific embodiment of this application, the pre-trained convolutional neural network is implemented using a YOLO convolutional neural network. The YOLO convolutional neural network boasts high detection speed and good accuracy, enabling rapid localization and classification of targets in images. During feature extraction, the YOLO convolutional neural network utilizes its unique network structure, such as the Darknet backbone network and feature pyramid network, to perform multi-level feature mining on the third image. Shallow convolutional layers primarily capture low-level features such as edges and corners, which are crucial for initially locating potential target regions. Mid-level convolutional layers further extract intermediate features such as texture and shape. Deep convolutional layers learn high-level semantic information from the image. Through this multi-level feature extraction method, the YOLO convolutional neural network can acquire rich multi-visual morphological details from the third image and transform them into a set of feature vectors. These feature vector sets not only contain the basic features of the target to be detected but also encompass feature information at different scales and levels. Meanwhile, the YOLO convolutional neural network employs data augmentation techniques during training, such as random pruning, rotation, and flipping, which increases the diversity of training data, improves the model's generalization ability, and enables it to perform better when facing object detection tasks in different scenarios.
[0028] S205, use the feature vector set as the data training set, and use the data training set to train the preset initial classification model to obtain the object detection model; Specifically, the feature vector set obtained through feature extraction is first organized and labeled according to the target type, location, shape, or sample collection conditions to construct a structured training set. Next, the training set is input into a pre-defined classification model, which can be a lightweight neural network, a deep fully connected network, a Softmax classifier, a support vector machine (SVM), a gradient boosting classifier, etc. During model training, the system performs forward propagation to calculate the error between the output class and the true class, and evaluates the fit of the model's current parameters based on a loss function (such as cross-entropy loss, contrastive loss, etc.). Subsequently, gradient updates are performed on the model parameters to progressively optimize the classification model. Furthermore, a validation set can be added during training iterations for model performance monitoring. By comparing the training error with the validation error, hyperparameters such as the learning rate, regularization parameters, and network depth are adjusted to prevent overfitting and improve the model's generalization ability to different target classes. After multiple rounds of training and parameter updates, a target detection model capable of accurately distinguishing target class features is finally obtained.
[0029] Furthermore, after training with the training dataset, a corresponding weight file is obtained. This weight file is then used to perform validation inference on the third image, producing a result output image. By comparing the result output image with a manually labeled standard image, key metrics such as accuracy, recall, and F1 score are calculated. If the model performance does not reach a preset threshold, the system can return to the feature extraction stage to re-optimize the network structure or adjust training parameters. Alternatively, data augmentation techniques can be used to expand the diversity of training samples, such as rotating, scaling, cropping, or adding noise to the original image, to improve the model's adaptability to complex processing scenarios. During the validation inference process, the system can also record the model's recognition time and false detection rate for different types of defects (such as cracks, dents, burrs, and dimensional deviations).
[0030] During training iterations, an early stopping mechanism is added. For example, after 50,000 training rounds, if the evaluation metric does not improve after 500 training rounds, training is stopped to conserve computing resources and prevent the continued aggravation of overfitting. Simultaneously, to improve model training efficiency, a distributed training framework can be used, distributing training tasks across multiple computing devices for parallel processing, or mixed-precision training techniques can be used to reduce computational resource consumption and accelerate model convergence. After training, the object detection model is lightweighted. This can be achieved through model pruning to remove redundant neurons and connections, quantization compression to convert weight parameters from high-precision floating-point to low-precision integer types, or knowledge distillation techniques to transfer knowledge from complex models to lightweight models, enabling rapid deployment and real-time detection on terminal devices.
[0031] S206: Receive the detection instruction, acquire the image to be detected, and perform target detection on the image to be detected based on the target detection model to obtain the target detection result.
[0032] Specifically, firstly, when the system receives a detection command from the operating terminal, automated control system, or detection process scheduling module, it calls the image acquisition device to acquire the image to be detected of the current workpiece or processing process. To ensure the stability of the detection, the camera parameters can be automatically calibrated before acquisition, including resetting the exposure, gain, and focal length to adapt to changes in lighting and position at the current processing site. The acquired image to be detected is then input into the trained target detection model, which performs feature extraction, feature comparison, and classification inference processes on the image to identify potential target regions within the image.
[0033] The solution presented in this application is applied to logo inspection on processed products. In logo inspection scenarios, this solution is adaptable to the marking inspection needs of various material surfaces, including metal, plastic, and glass. For metal logos, it eliminates reflective interference and identifies the integrity of the logo outline; for plastic logos, it uses a multi-angle lighting scheme to capture surface printing defects, such as uneven ink distribution and missing characters; in glass inspection, it checks whether the logo etching depth meets standards. The system also supports a dynamic inspection mode, enabling real-time tracking and inspection of moving workpieces on the production line to ensure that the logo on each product meets quality requirements.
[0034] Furthermore, the step of denoising the first image to obtain the second image, taking into account changes in lighting and background interference in the processing environment, specifically includes: Illumination intensity is estimated in the first image, and the image brightness is normalized based on the illumination intensity distribution to obtain the first denoised image; The first denoised image is processed using a multi-scale filtering method to separate random noise from background texture interference, and noise features are extracted. Based on noise features, spatial domain smoothing is performed on the first denoised image to obtain the second denoised image, wherein spatial domain smoothing is used to suppress random noise and preserve the image edge structure. The second denoised image is subjected to frequency domain filtering to obtain the third denoised image; The third denoised image is then subjected to pixel reconstruction and image detail correction to obtain the second image.
[0035] Specifically, when estimating the illumination intensity of the first image, the system utilizes the pixel grayscale value distribution information of the image, combined with a preset illumination model, to analyze the illumination intensity in different regions of the image. Based on the illumination intensity distribution results, the brightness of the image is adjusted to a uniform range to eliminate the impact of uneven illumination on image quality, thereby obtaining the first denoised image.
[0036] When applying multi-scale filtering to the first denoised image, the system uses filters of different scales, such as Gaussian filters and wavelet filters, to filter the image. In this way, random noise in the image can be separated from background texture interference, and noise features can be extracted.
[0037] Based on noise characteristics, when performing spatial domain smoothing on the first denoised image, the image is smoothed, which can suppress random noise while maintaining the edge structure of the image, thus obtaining the second denoised image.
[0038] When applying frequency domain filtering to the second denoised image, the system converts the image from the spatial domain to the frequency domain and uses a frequency domain filter to filter the image. Frequency domain filtering can effectively remove periodic noise or interference at specific frequencies from the image, thus obtaining the third denoised image.
[0039] When performing pixel-level reconstruction and image detail correction on the third denoised image, pixel-level reconstruction and detail correction are performed. These operations can recover detail information that may have been lost during the denoising process, improve the image's clarity and quality, and ultimately obtain the second image.
[0040] Further, the step of enhancing the contrast of the second image to obtain the third image specifically includes: Perform local brightness statistical analysis on the second image to obtain the brightness distribution characteristics of the second image in different regions; The grayscale range of the second image is dynamically expanded based on the brightness distribution characteristics to obtain the first enhanced image; The first enhanced image is subjected to local region histogram equalization to obtain the second enhanced image, wherein the histogram equalization process is used to further enhance the edge and texture details of the potential target region; The second enhanced image is then smoothed and pixel-matched to obtain the third image.
[0041] Specifically, when performing local brightness statistical analysis on the second image, the system divides the second image into multiple local regions according to a preset image segmentation strategy, and calculates the mean and variance of brightness for each region to obtain the brightness distribution characteristics of the second image in different regions. These statistics can reflect the brightness differences in different regions of the image.
[0042] Based on the brightness distribution characteristics, the system dynamically adjusts the grayscale range of each local area according to its brightness distribution. For example, for areas with low brightness, the grayscale range is appropriately stretched to increase brightness; for areas with high brightness, the grayscale range may be compressed to avoid overexposure. Through this dynamic expansion, a first enhanced image is obtained, initially improving the overall contrast of the image.
[0043] When performing local region histogram equalization on the first enhanced image, the system redistributes the grayscale values of pixels in each local region, making the pixel distribution more uniform across the grayscale range. This further enhances the edges and texture details of potential target regions, as these details typically exhibit variations in grayscale, and histogram equalization makes these variations more pronounced, thus producing the second enhanced image.
[0044] When performing smoothing correction and pixel coordination on the second enhanced image, the system eliminates noise and artifacts that may be introduced during histogram equalization. Simultaneously, to maintain the overall visual quality of the image, the system also performs pixel coordination operations, adjusting the grayscale relationships between adjacent pixels to make the transition between enhanced and unenhanced areas more natural, ultimately resulting in the third image.
[0045] Further, please refer to Figure 3 The steps include: obtaining pixel distribution information from a third image; extracting multi-layer features from the third image using a pre-trained convolutional neural network based on the pixel distribution information; and obtaining a set of feature vectors containing multiple visual morphological details. S301, Perform pixel distribution statistics on the third image to generate a pixel density map to guide the activation of the convolutional layer; S302, input the third image and pixel density map into the pre-trained convolutional neural network, and adjust the initial receptive field of the convolutional neural network according to the pixel density map to activate the convolutional neural network. S303, after activating the convolutional neural network, performs layer-by-layer feature extraction on the third image through the multi-scale convolutional layer, residual module and feature fusion layer of the convolutional neural network to obtain an intermediate feature set containing target contour, texture and edge information; S304 performs feature concatenation and vectorization on the intermediate feature set to form a feature vector set containing multiple visual morphological details.
[0046] Specifically, when performing pixel distribution statistics on the third image, the system iterates through each pixel in the image, counts the frequency of its pixel value, and generates a pixel density map accordingly. This map can intuitively reflect the distribution of different pixel values in the image.
[0047] The third image and the generated pixel density map are input together into a pre-trained convolutional neural network. During the input process, the system dynamically adjusts the initial receptive field size of the convolutional neural network based on the pixel density map. The receptive field is the image region "observed" by each neuron in the convolutional neural network. By adjusting the receptive field, the network can focus more on areas with dense pixel distribution in the image, thereby more effectively activating the convolutional neural network and improving the accuracy of feature extraction.
[0048] After activating the convolutional neural network, the network utilizes its multi-scale convolutional layers, residual modules, and feature fusion layers to extract features from the third image layer by layer. Multi-scale convolutional layers can capture feature information at different scales in the image, effectively extracting everything from subtle edges to large texture structures. The residual modules, by introducing skip connections, solve the gradient vanishing problem during deep network training, enabling the network to learn deeper features. The feature fusion layer combines features from different levels to form a richer and more comprehensive intermediate feature set, which contains key information such as the target's contours, textures, and edges.
[0049] The extracted intermediate feature set undergoes feature concatenation and vectorization. Feature concatenation combines features from different levels and scales to form a more complete feature representation; vectorization converts the concatenated features into vector form. Through this step, a feature vector set containing multiple visual morphological details is ultimately formed.
[0050] Furthermore, the step of extracting features from the third image layer by layer through multi-scale convolutional layers, residual modules, and feature fusion layers of a convolutional neural network to obtain an intermediate feature set containing target contour, texture, and edge information specifically includes: The third image is sequentially fed into a multi-scale convolutional layer, and the basic edge features and texture details of the third image are extracted through the multi-scale convolutional kernel. The output of the multi-scale convolutional layer is input into the residual module, and the deep features are enhanced through the residual connection structure to obtain the preliminary target morphology features; By using multi-scale convolutional layers to calculate the feature response at different scales of the initial target morphology features, the contour, texture, and edge structure features of targets of different sizes are obtained. The contour, texture, and edge structure features of targets of different sizes are input into the feature fusion layer. The features at each scale are integrated and enhanced through a preset spatial attention mechanism to obtain the target fusion feature map. The target fusion feature map is dimensionally normalized and uniformly encoded to output an intermediate feature set containing target contour, texture and edge information.
[0051] Specifically, when the third image is input into the multi-scale convolutional layer, the system uses convolutional kernels of different sizes to perform convolution operations on the image. These convolutional kernels can capture the basic edge features and texture details of the image at different scales. For example, small-sized convolutional kernels can extract subtle edges in the image, while large-sized convolutional kernels can capture larger texture structures. In this way, the multi-scale convolutional layer can comprehensively extract the basic feature information in the image.
[0052] The outputs of multi-scale convolutional layers are fed into the residual module. The residual module introduces skip connections to add the input features to the features obtained after multiple convolutions. This structure effectively solves the vanishing gradient problem during deep network training, enabling the network to learn deeper feature information. Through the processing of the residual module, the system obtains preliminary target morphological features, which include the approximate outline and texture information of the target to be detected.
[0053] The system then utilizes multi-scale convolutional layers to calculate feature responses at different scales for the initial target morphology features. Different sized convolutional kernels are used to further convolve the feature map, capturing the contour, texture, and edge structure features of targets of varying sizes. For example, for smaller targets, smaller kernels can more accurately capture their contours and textures; while for larger targets, larger kernels can better extract their overall structural features. Through this step, the system obtains more comprehensive and detailed feature information about the target to be detected.
[0054] The contour, texture, and edge structure features of targets of different sizes are input into the feature fusion layer. The feature fusion layer employs a pre-defined spatial attention mechanism to integrate and enhance features at each scale. This spatial attention mechanism dynamically adjusts the weights of different regions in the feature map based on their importance, making the network focus more on the target's location and thus improving the accuracy of feature extraction. Through the processing of the feature fusion layer, the system obtains a target fusion feature map, which integrates target feature information from different scales, more comprehensively reflecting the morphological details of the target to be detected.
[0055] The system performs dimensionality normalization and unified encoding on the target fusion feature map. It normalizes the feature map to ensure a uniform dimensionality. Simultaneously, through unified encoding, it converts the feature map into vector form, creating an intermediate feature set containing target contour, texture, and edge information.
[0056] Furthermore, the output of the multi-scale convolutional layer is input into the residual module to enhance deep features through the residual connection structure, thereby obtaining preliminary target morphological features. This process specifically includes: The output of the multi-scale convolutional layer is input into the main branch and shortcut branch of the residual module; The features output by the main branch and the features output by the shortcut branch are superimposed element by element to obtain the residual enhancement features; The residual enhancement features are processed by activation functions, which further enhance the expression of deep features and improve the discernibility of details in multiple visual forms. The activated residual enhancement features are used as preliminary target morphological features, and the preliminary target morphological features are output.
[0057] Specifically, when the outputs of the multi-scale convolutional layers are input into the main branch and the shortcut branch of the residual module, the main branch will undergo a series of operations such as convolution and batch normalization to further extract and transform features, while the shortcut branch will directly pass the input features without performing complex transformations. This design allows the network to retain more original information during training.
[0058] The features output by the main branch and the shortcut branch have the same dimensionality. The system then performs element-wise stacking of these two branch outputs. Element-wise stacking is a simple and effective feature fusion method that combines the different features learned by the two branches to obtain a richer and more comprehensive feature representation. Through this step, the system obtains residual enhanced features, which contain both the deep feature information learned by the main branch and the original feature information passed from the shortcut branch.
[0059] When applying activation functions to residual augmentation features, the system employs activation functions such as ReLU (Rectified LinearUnit). The role of activation functions is to introduce non-linearity into the network, enabling it to learn more complex and abstract feature representations. Through activation function processing, the deep feature representation of residual augmentation features is further enhanced, while the discernibility of multi-visual morphological details is also improved. This means the network can more accurately capture the target visual features in the image.
[0060] The activated residual enhancement features are used as preliminary target morphological features, and these features are output. These preliminary target morphological features already contain key information such as the approximate outline and texture of the target to be detected. The system will then pass these features to the next layer of the network for further processing and analysis to achieve a more accurate and efficient visual detection task.
[0061] Furthermore, the steps of using the feature vector set as the data training set and training the preset initial classification model using the data training set to obtain the object detection model specifically include: The feature vector set is labeled and classified according to the type of the target to be detected, forming a structured data training set; The structured training data set is input into the preset initial classification model, and forward propagation and loss calculation are performed on the model parameters based on the feature labels; The parameters of the initial classification model are iteratively updated based on the loss calculation results to improve the model's ability to distinguish different target features; During the iterative update process, the initial classification model is evaluated using a validation set and its hyperparameters are adjusted to obtain a convergent and stable object detection model.
[0062] Specifically, taking logo detection as an example, the system manually or automatically labels the image region corresponding to each vector in the feature vector set, clarifying whether it contains a logo, the specific type of the logo, and the approximate location of the logo in the image. Then, based on this labeling information, the feature vector set is divided into different category groups, forming a structured data training set. The data training set has clear category labels, which facilitates model learning and training.
[0063] The structured training dataset is input into a pre-defined initial classification model, which can be a support vector machine, random forest, or deep neural network. Upon receiving the input feature vectors and their corresponding feature labels, the model performs a forward propagation process based on the pre-defined network structure and initial parameters. During forward propagation, the feature vectors are calculated and transformed through each layer of the network, ultimately outputting a prediction result. The system compares this prediction result with the true feature labels and calculates the difference between them using a pre-defined loss function (such as cross-entropy loss).
[0064] Based on the loss calculation results, the system employs optimization algorithms such as gradient descent to iteratively update the parameters of the initial classification model. Specifically, the system calculates the partial derivatives (gradients) of the loss function with respect to each model parameter, and then adjusts the parameter values along the gradient descent direction to reduce the loss value. By continuously repeating the process of forward propagation, loss calculation, and parameter update, the model can gradually learn the differences between features of different target categories, thereby improving its ability to distinguish between different target features. For example, for logos of different shapes, the model can learn subtle differences in feature vectors to more accurately distinguish between logos.
[0065] During iterative updates, to avoid overfitting and ensure the model's generalization ability, the system divides the training data into a training set and a validation set. After each or several rounds of parameter updates, the system uses the validation set to evaluate the current model, calculating metrics such as accuracy, precision, and recall. Through this validation set evaluation and hyperparameter tuning process, the model performance is continuously optimized until its performance on both the training and validation sets reaches a stable state, and the loss value no longer decreases significantly. The resulting model is a convergent and stable object detection model, capable of accurately detecting and classifying objects in the input image.
[0066] In one specific embodiment of this application, the initial classification model can be implemented using a Support Vector Machine (SVM). For example, when detecting the logo of a processed workpiece, the initial classification model uses an SVM to construct a multi-classifier structure. Specifically, the system constructs multiple binary SVM sub-models based on different logo target types (such as font deformation, color deviation, pattern incompleteness, etc.). Each sub-model maps the input features to a high-dimensional space through a kernel function, searching for the optimal classification hyperplane to distinguish the specific target type from normal samples. During the training phase, the system uses a one-vs-one strategy to construct a multi-classification architecture, training one sub-classifier, and finally determining the final target type by combining the prediction results of all sub-models through a voting mechanism. To improve model robustness, the system introduces slack variables and regularization parameters, and optimizes the penalty coefficient C and kernel function parameter γ through cross-validation, so that the model achieves a balance between training error and generalization ability. Furthermore, to address the potential class imbalance issue in workpiece logo detection, the system employs a weighted SVM method. This dynamically adjusts the margin weights of the classification hyperplane based on the number of samples in each class, ensuring the model maintains a high recognition rate even for minority class targets. At the feature input level, the system normalizes the feature vector set extracted in the preceding steps, which includes target contour, texture, and edge information. This eliminates dimensional differences between different feature dimensions, enabling the SVM model to make classification decisions based on a unified feature scale. Through this structured modeling approach, the initial classification model fully leverages the advantages of support vector machines in small-sample learning. Simultaneously, through multi-classification strategies and parameter optimization mechanisms, it achieves high-precision and high-reliability detection of processed workpiece logos.
[0067] In one specific embodiment of this application, please refer to Figure 4 The visual inspection method is applied to the logo inspection of display devices in the production process. The display device to be inspected is placed on the inspection table, which is equipped with an electrical cabinet, inspection station, equipment mounting rack, camera and light source. The electrical cabinet contains an industrial control computer, power supply, monitor, etc. The industrial control computer is equipped with an MES server and a vision server. The MES server is used to store standard logos and various image sets, and the vision server is used to execute the implementation steps of the above-mentioned visual inspection method. The camera and light source are respectively set on the equipment mounting rack. The display device to be inspected is placed at the inspection station by the robotic arm of the production line and the information of the display device to be inspected is confirmed by the barcode scanner.
[0068] After information confirmation, the logo inspection process begins. Specifically, the vision server controls the light source to uniformly illuminate the logo area of the display device to be inspected according to preset lighting parameters (such as brightness, color temperature, and illumination angle) to eliminate ambient light interference and highlight the logo's visual features. Subsequently, the camera is controlled to acquire images of the logo area of the display device at the inspection station. A high-resolution industrial camera can be used, and shooting parameters such as focal length and exposure time are adjusted according to the logo's size and position to ensure that the acquired logo image is clear, detailed, and that the logo area occupies an appropriate proportion in the image. The acquired image data is transmitted to the vision server via a data interface. The vision server performs preprocessing on the original image, such as image denoising, contrast enhancement, and distortion correction, to optimize image quality. After preprocessing, the vision server follows the aforementioned visual inspection method steps, sequentially performing feature response calculations for multi-scale convolutional layers, preliminary target morphology feature extraction from the residual module, spatial attention mechanism integration in the feature fusion layer, dimensionality normalization and unified encoding, and classification and recognition of the processed features using a trained target detection model. Ultimately, it determines whether the logo on the display device has defects such as font distortion, color deviation, or pattern incompleteness, and feeds the inspection results (e.g., pass / fail, specific defect type, etc.) back to the MES server. Simultaneously, the results can be displayed in real-time on the monitor so that operators can promptly understand the inspection status. If a non-conforming product is detected, the system can further trigger subsequent sorting or alarm mechanisms, achieving automated and high-precision inspection of the logo quality on display devices during the production process.
[0069] Furthermore, for logos that fail the initial inspection, manual review is conducted to ensure accuracy and prevent misjudgments due to complex circumstances. Manual reviewers will consult the system's stored images of failed logos and their corresponding feature analysis reports, combining their professional experience to verify the type and severity of the defects. If the manual review determines the logo is acceptable, the result is sent to the MES server to correct the original inspection result, and the sample is added to the abnormal sample library in the data training set for subsequent optimization and iteration of the target detection model. If the manual review determines the logo is unacceptable, the product is categorized based on the defect type. For example, minor, repairable color deviation defects are routed to the repair station for rework; severe, irreparable defects such as font distortion or pattern incompleteness are directly deemed scrap and enter the disposal process.
[0070] In addition, for logos that fail the test, the system will also push an alarm message to WeChat Work, transmitting the corresponding image of the non-compliant logo, the product's serial number, time, and other information to WeChat Work.
[0071] The dual verification mechanism of "automated detection + manual supplementary inspection" not only ensures the efficiency of detection, but also minimizes the rate of missed detection and false detection, thereby effectively improving the reliability and rigor of logo quality control in the production process of display devices.
[0072] In this embodiment, the visual inspection method operates on an electronic device (e.g., Figure 1 The server shown can receive instructions or acquire data via wired or wireless connection. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra-wideband) connections, and other currently known or future wireless connection methods.
[0073] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).
[0074] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0075] Further reference Figure 5 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of a visual inspection device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0076] like Figure 5 As shown, the visual inspection device 400 described in this embodiment includes: The real-time imaging module 501 is used to acquire raw image data captured in real time from the processing environment to obtain the first image; The noise reduction module 502 is used to perform noise reduction processing on the first image to obtain the second image, taking into account changes in lighting and background interference in the processing environment. The contrast enhancement module 503 is used to enhance the contrast of the second image to obtain the third image, wherein the potential target visual form in the image is highlighted by the contrast enhancement. The feature extraction module 504 is used to obtain the pixel distribution information of the third image. Based on the pixel distribution information, a pre-trained convolutional neural network is used to extract multi-layer features from the third image to obtain a set of feature vectors containing multiple visual morphological details. The classification training module 505 is used to use the feature vector set as the data training set and to train the preset initial classification model using the data training set to obtain the object detection model. The target detection module 506 is used to receive detection instructions, acquire the image to be detected, and perform target detection on the image to be detected based on the target detection model to obtain the target detection result.
[0077] Furthermore, the noise reduction processing module 502 specifically includes: The brightness normalization submodule is used to estimate the illumination intensity of the first image, and normalize the image brightness based on the illumination intensity distribution result to obtain the first denoised image. The multi-scale filtering submodule is used to separate random noise and background texture interference from the first denoised image using a multi-scale filtering method, and to extract noise features; The spatial domain smoothing submodule is used to perform spatial domain smoothing on the first denoised image based on noise features to obtain the second denoised image, wherein spatial domain smoothing is used to suppress random noise and preserve the image edge structure. The frequency domain filtering submodule is used to apply frequency domain filtering to the second denoised image to obtain the third denoised image. The image reconstruction submodule is used to perform pixel reconstruction and image detail correction on the third denoised image to obtain the second image.
[0078] Furthermore, the contrast enhancement module 503 specifically includes: The brightness statistics submodule is used to perform local brightness statistical analysis on the second image and obtain the brightness distribution characteristics of the second image in different regions. The contrast stretching submodule is used to dynamically expand the grayscale range of the second image based on the brightness distribution characteristics to obtain the first enhanced image; The histogram equalization submodule is used to perform local region histogram equalization processing on the first enhanced image to obtain the second enhanced image, wherein the histogram equalization processing is used to further enhance the edge and texture details of the potential target region. The smoothing correction submodule is used to perform smoothing correction and pixel coordination on the second enhanced image to obtain the third image.
[0079] Further, please refer to Figure 6 The feature extraction module 504 specifically includes: The pixel statistics submodule 601 is used to perform pixel distribution statistics on the third image and generate a pixel density map to guide the activation of the convolutional layer. The model activation submodule 602 is used to input the third image and pixel density map into the pre-trained convolutional neural network and adjust the initial receptive field of the convolutional neural network according to the pixel density map to activate the convolutional neural network. The feature extraction submodule 603 is used to extract features from the third image layer by layer after activating the convolutional neural network through the multi-scale convolutional layer, residual module and feature fusion layer of the convolutional neural network to obtain an intermediate feature set containing target contour, texture and edge information. The feature splicing submodule 604 is used to splice and vectorize the intermediate feature set to form a feature vector set containing multiple visual morphological details.
[0080] Furthermore, the feature extraction submodule 603 specifically includes: The feature convolution unit is used to sequentially input the third image into the multi-scale convolutional layer, and extract the basic edge features and texture details of the third image through the multi-scale convolutional kernel; The feature enhancement unit is used to input the output of the multi-scale convolutional layer into the residual module, and enhance the deep features through the residual connection structure to obtain the preliminary target morphology features; The feature response unit is used to calculate the feature response at different scales on the initial target morphological features through multi-scale convolutional layers, so as to obtain the contour, texture and edge structure features of targets of different sizes; The feature fusion unit is used to input the contour, texture, and edge structure features of targets of different sizes into the feature fusion layer. Through a preset spatial attention mechanism, the features at each scale are integrated and enhanced to obtain the target fusion feature map. The dimension normalization unit is used to normalize and uniformly encode the target fused feature map, and outputs an intermediate feature set containing target contour, texture and edge information.
[0081] Furthermore, the feature enhancement unit specifically includes: The feature input subunit is used to input the output of the multi-scale convolutional layer into the main branch and shortcut branch of the residual module; The element-overlay subunit is used to overlay the features output by the main branch with the features output by the shortcut branch element by element to obtain residual enhancement features. The activation processing subunit is used to perform activation function processing on the residual enhancement features. The activation function processing further enhances the expression of deep features and improves the discernibility of multi-visual morphological details. The residual enhancement output subunit is used to take the activated residual enhancement features as preliminary target morphological features and output the preliminary target morphological features.
[0082] Furthermore, the classification training module 505 specifically includes: The classification and labeling submodule is used to label and classify the feature vector set according to the type of the target to be detected, forming a structured data training set; The loss calculation submodule is used to input the structured data training set into the preset initial classification model, and perform forward propagation and loss calculation on the model parameters based on the feature labels; The iterative update submodule is used to iteratively update the parameters of the initial classification model based on the loss calculation results, so as to improve the model's ability to distinguish different target features; The evaluation and adjustment submodule is used to evaluate the initial classification model on a validation set and adjust its hyperparameters during the iterative update process, so as to obtain a convergent and stable object detection model.
[0083] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 7 , Figure 7 This is a basic structural block diagram of the computer device in this embodiment.
[0084] The computer device includes a memory and a processor. The memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the visual detection method described above.
[0085] The computer device 7 includes a memory 71, a processor 72, and a network interface 73 that are interconnected via a system bus. It should be noted that only a computer device 7 with a memory 71, a processor 72, and a network interface 73 is shown in the figure; however, it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0086] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.
[0087] The memory 71 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 71 may be an internal storage unit of the computer device 7, such as the hard disk or memory of the computer device 7. In other embodiments, the memory 71 may also be an external storage device of the computer device 7, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 7. Of course, the memory 71 may include both the internal storage unit and its external storage device of the computer device 7. In this embodiment, the memory 71 is typically used to store the operating system and various application software installed on the computer device 7, such as computer-readable instructions for visual inspection methods. In addition, the memory 71 can also be used to temporarily store various types of data that have been output or will be output.
[0088] In some embodiments, the processor 72 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 72 is typically used to control the overall operation of the computer device 7. In this embodiment, the processor 72 is used to execute computer-readable instructions stored in the memory 71 or to process data, for example, to execute computer-readable instructions for the visual inspection method.
[0089] The network interface 73 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 7 and other electronic devices.
[0090] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the visual inspection method described above.
[0091] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0092] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.
Claims
1. A visual inspection method, characterized in that, include: The first image is obtained by acquiring raw image data captured in real time from the processing environment. To address the lighting changes and background interference in the processing environment, the first image is denoised to obtain the second image; The second image is contrast-enhanced to obtain the third image; The pixel distribution information of the third image is obtained, and a pre-trained convolutional neural network is used to extract multi-layer features from the third image based on the pixel distribution information to obtain a feature vector set containing multiple visual morphological details. The feature vector set is used as the data training set, and the preset initial classification model is trained using the data training set to obtain the target detection model; The system receives a detection instruction, acquires the image to be detected, and performs target detection on the image based on the target detection model to obtain the target detection result.
2. The visual inspection method as described in claim 1, characterized in that, The step of denoising the first image to obtain the second image in response to changes in lighting and background interference in the processing environment specifically includes: Illumination intensity is estimated in the first image, and the image brightness is normalized based on the illumination intensity distribution to obtain the first denoised image; The first denoised image is processed using a multi-scale filtering method to separate random noise from background texture interference, and noise features are extracted. Based on the noise characteristics, spatial domain smoothing is performed on the first denoised image to obtain the second denoised image; The second denoised image is subjected to frequency domain filtering to obtain the third denoised image; The third denoised image is then subjected to pixel reconstruction and image detail correction to obtain the second image.
3. The visual inspection method as described in claim 1, characterized in that, The step of enhancing the contrast of the second image to obtain the third image specifically includes: Perform local brightness statistical analysis on the second image to obtain the brightness distribution characteristics of the second image in different regions; Based on the brightness distribution characteristics, the grayscale range of the second image is dynamically expanded to obtain the first enhanced image; Perform local region histogram equalization on the first enhanced image to obtain the second enhanced image; The second enhanced image is then smoothed and pixel-coordinated to obtain the third image.
4. The visual inspection method as described in claim 1, characterized in that, The step of obtaining the pixel distribution information of the third image, and extracting multi-layer features from the third image using a pre-trained convolutional neural network based on the pixel distribution information to obtain a feature vector set containing multiple visual morphological details, specifically includes: Perform pixel distribution statistics on the third image to generate a pixel density map to guide the activation of the convolutional layer; The third image and the pixel density map are input into a pre-trained convolutional neural network, and the initial receptive field of the convolutional neural network is adjusted according to the pixel density map to activate the convolutional neural network. After activating the convolutional neural network, the third image is subjected to layer-by-layer feature extraction through the multi-scale convolutional layer, residual module and feature fusion layer of the convolutional neural network to obtain an intermediate feature set containing target contour, texture and edge information; The intermediate feature set is then processed by feature concatenation and vectorization to form a feature vector set containing multiple visual morphological details.
5. The visual inspection method as described in claim 4, characterized in that, The step of extracting features from the third image layer by layer through the multi-scale convolutional layers, residual modules, and feature fusion layers of the convolutional neural network to obtain an intermediate feature set containing target contour, texture, and edge information specifically includes: The third image is sequentially input into the multi-scale convolutional layer, and the basic edge features and texture details of the third image are extracted through the multi-scale convolutional kernel; The output of the multi-scale convolutional layer is input to the residual module, and the deep features are enhanced through the residual connection structure to obtain the preliminary target morphology features; The multi-scale convolutional layer is used to calculate the feature response of the preliminary target morphology features at different scales to obtain the contour, texture, and edge structure features of targets of different sizes. The contour, texture, and edge structure features of the targets of different sizes are input into the feature fusion layer. The features at each scale are integrated and enhanced through a preset spatial attention mechanism to obtain a target fusion feature map. The target fusion feature map is dimensionally normalized and uniformly encoded to output the intermediate feature set containing target contour, texture and edge information.
6. The visual inspection method as described in claim 5, characterized in that, The step of inputting the output of the multi-scale convolutional layer into the residual module, and enhancing deep features through the residual connection structure to obtain preliminary target morphological features, specifically includes: The output of the multi-scale convolutional layer is input into the main branch and shortcut branch of the residual module; The features output by the main branch and the features output by the shortcut branch are superimposed element by element to obtain the residual enhancement features; The residual enhancement features are processed by activation functions to further enhance the expression of deep features and improve the discernibility of details in multiple visual forms; The activated residual enhancement feature is used as the preliminary target morphological feature, and the preliminary target morphological feature is output.
7. The visual inspection method as described in claim 1, characterized in that, The step of using the feature vector set as a training set and training a preset initial classification model using the training set to obtain an object detection model specifically includes: The feature vector set is labeled and classified according to the type of the target to be detected, forming a structured data training set; The structured data training set is input into a preset initial classification model, and forward propagation and loss calculation are performed on the model parameters based on the feature labels; The parameters of the initial classification model are iteratively updated based on the loss calculation results to improve the model's ability to distinguish different target features; During the iterative update process, the initial classification model is evaluated using a validation set and its hyperparameters are adjusted to obtain a convergent and stable target detection model.
8. A visual inspection device, characterized in that, include: The real-time imaging module is used to acquire raw image data captured in real time from the processing environment to obtain the first image; The noise reduction module is used to perform noise reduction processing on the first image to obtain a second image, taking into account changes in lighting and background interference in the processing environment. A contrast enhancement module is used to enhance the contrast of the second image to obtain a third image; The feature extraction module is used to obtain the pixel distribution information of the third image, and extract multi-layer features from the third image using a pre-trained convolutional neural network based on the pixel distribution information to obtain a feature vector set containing multiple visual morphological details. The classification training module is used to use the feature vector set as a data training set and to train a preset initial classification model using the data training set to obtain an object detection model. The target detection module is used to receive detection instructions, acquire the image to be detected, and perform target detection on the image to be detected based on the target detection model to obtain the target detection result.
9. A computer device, characterized in that, The device includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the visual inspection method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the visual inspection method as described in any one of claims 1 to 7.