Robot based on AI visual quality inspection system
By integrating an AI visual quality inspection system with industrial cameras and thermal imaging cameras, combined with a deep learning processor and wireless communication module, the problems of blind spots and high misjudgment rates in defect detection in complex environments are solved, and efficient and real-time defect identification and feedback are achieved.
Patent Information
- Application Number
- CN202510698325.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-19
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing industrial product defect detection technology lacks robustness in complex environments, making it difficult to achieve comprehensive defect identification. It also has low detection efficiency and poor adaptability. In particular, there are problems of blind spots and high misjudgment rates in the inspection of large workpieces, outdoor structures and special-shaped components.
It uses a robot based on an AI visual quality inspection system, integrated with industrial cameras and thermal imaging cameras, processes image data through a multimodal fusion module, combines with a deep learning processor for defect identification, and transmits the results in real time through a wireless communication module. It is equipped with a power management unit to ensure stable power supply.
It achieves high-precision defect identification and real-time feedback in complex environments, improves the integrity and efficiency of detection, solves the problems of blind spots and high misjudgment rates in detection, and supports flexible deployment and remote feedback.
Smart Images

Figure CN120673129A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robotics technology, and in particular to a robot based on an AI visual quality inspection system. Background Art
[0002] In the current technological landscape, quality inspection of industrial products, particularly surface defect detection, generally relies on fixed inspection equipment. This type of inspection typically requires products to be transported to a specific inspection location during the production process, where statically mounted cameras or sensors capture and process images. While this inspection model has some applicability in closed, clearly structured assembly line scenarios, it lacks versatility, flexibility, and adaptability, making it difficult to cover industrial sites with widespread distribution, complex structures, or high mobility.
[0003] Especially when faced with inspection tasks such as large workpieces, outdoor structures, and special-shaped components, traditional static visual inspection equipment is often limited by fixed shooting angles, limited field of view, and significant blind spots, making it difficult to guarantee comprehensive and accurate defect identification. More importantly, under various non-ideal conditions such as large variations in ambient light, complex target surface materials, and abnormal heat distribution, traditional single-modality inspection methods lack robustness and are easily affected by background noise, artifacts, and occlusions. This leads to high misjudgment rates and frequent missed detections, severely restricting the intelligence level and industrial adaptability of quality inspection systems.
[0004] Despite significant progress in recent years in robotic platforms' mobility control and adaptability to specific scenarios—tracked robots, wheeled mobile platforms, and even bionic quadruped robots are increasingly being used for inspection—defect detection capabilities still rely primarily on conventional cameras capturing images, followed by manual or weak algorithmic processing. Highly integrated systems with end-to-end intelligent recognition, positioning, and feedback capabilities have yet to emerge. Existing robotic quality inspection systems, in particular, suffer from low integration, delayed response times, and inefficient operations when it comes to fusion of multiple sensor data, multimodal image understanding, and real-time transmission of recognition results.
[0005] Therefore, the present invention proposes a robot based on an AI visual quality inspection system to address the deficiencies of the prior art. Summary of the Invention
[0006] The purpose of the present invention is to provide a robot based on an AI visual quality inspection system, which solves the problems of incomplete defect recognition, poor environmental adaptability and low detection efficiency in the prior art.
[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions: A robot based on an AI visual quality inspection system includes: A robot body, the robot body including a camera bracket having a movable structure and an adjustable angle; At least one set of industrial cameras and thermal imaging camera modules, mounted on the camera bracket, for synchronously collecting image data after the robot moves to the target detection area, the image data including visible light image data and thermal imaging image data; an image preprocessing module, connected to the industrial cameras and thermal imaging camera modules, for preprocessing the collected image data, the preprocessing including adaptive contrast enhancement, edge sharpening, and noise suppression of the collected image data; a multimodal fusion module, connected to the image preprocessing module, for performing spatial registration and feature fusion on the preprocessed image data to generate fused image data; a deep learning processor, connected to the multimodal fusion module, for performing defect recognition analysis on the fused image data and outputting a defect recognition result including defect type, location, and recognition confidence; A wireless communication module, connected to the deep learning processor, is used to transmit the defect recognition results to the background monitoring system in real time via wireless mode; A power management unit is connected to the robot body, the industrial camera and thermal imaging camera module, the image preprocessing module, the multimodal fusion module, the deep learning processor and the wireless communication module to provide stable working power.
[0008] Preferably, the mobile structure includes a mobile wheel chassis, which is installed at the bottom of the robot body; the camera bracket includes a rotating bearing base, which is installed at the top of the robot body; and a multi-degree-of-freedom robotic arm is installed on the rotating bearing base.
[0009] Preferably, the industrial camera and thermal imaging camera module are installed on top of a multi-degree-of-freedom robotic arm.
[0010] Preferably, the image preprocessing module includes: The adaptive contrast enhancement submodule is used to dynamically adjust the contrast of the collected image data so that the target area and the background area in the image data are clearly distinguished in terms of brightness level; The edge sharpening submodule is connected to the adaptive contrast enhancement submodule to further enhance the edge and contour features of objects in the image to highlight the boundaries of defects; The noise suppression submodule is connected to the edge sharpening submodule and is used to filter out random noise and unstructured interference signals in the image data.
[0011] Preferably, the multimodal fusion module includes: The spatial registration submodule is used to geometrically align the visible light image and the thermal imaging image in the preprocessed image data to ensure that the visible light image and the thermal imaging image correspond at the pixel level; The feature extraction submodule is connected to the spatial registration submodule and is used to extract key feature information of edges, textures and thermal distribution from the aligned visible light image and thermal imaging image respectively; The feature fusion submodule is connected to the feature extraction submodule and is used to fuse the key features of the visible light image and the thermal imaging image into multi-dimensional information to generate fused image data.
[0012] Preferably, the feature fusion submodule performs multi-dimensional information fusion on key features of the visible light image and the thermal imaging image using the following formula: F fusion =α·F visible +β·F thermal ; Among them, F visible Represents the feature vector of the visible light image; F thermal represents the eigenvector of the thermal imaging image, α and β are the weight coefficients of the visible light mode and the thermal imaging mode, respectively, and satisfy α+β=1; F fusion is the fused feature vector.
[0013] Preferably, the deep learning processor includes: The image encoding submodule is used to receive the fused image data and extract its spatial and texture features to generate a multi-scale feature representation. The defect recognition submodule is connected to the image encoding submodule and is used to identify defects in the image data based on the multi-scale feature representation using a pre-trained deep neural network model, and output the defect type and corresponding location coordinates. The confidence output submodule is connected to the defect recognition submodule and is used to extract and output the recognition confidence to form a complete defect recognition result.
[0014] Preferably, the confidence output submodule calculates the recognition confidence using the following formula: Among them, C represents the recognition confidence of the final output; P cls represents the probability of defect classification; P loc represents the positioning accuracy score; w1 and w2 are preset weight coefficients, and satisfy w1+w2=1; e is the base of the natural logarithm.
[0015] Preferably, the wireless communication module includes a Wi-Fi communication unit, a Bluetooth communication unit and a cellular communication unit; the wireless communication module transmits the defect identification result to the background monitoring system in real time via wireless means.
[0016] In summary, the present invention includes at least one of the following beneficial technical effects: 1. This invention utilizes a multimodal sensing architecture integrating an industrial camera and a thermal imaging camera, and employs a multimodal fusion module for spatial registration and feature fusion. This technology achieves the simultaneous identification of surface and internal defects in complex inspection environments. Compared to existing robotic inspection solutions that rely solely on a single-mode image source for defect identification, this approach addresses the limitations of existing robotic inspection solutions, such as insufficient perception of obscured and thermally anomaly-related defects and limited identification coverage, significantly improving the integrity and accuracy of robotic inspections.
[0017] 2. This invention employs a deep learning processing architecture consisting of image encoding, defect recognition, and confidence output. It utilizes a trained deep neural network model to extract features and identify defects from fused images, achieving the technical effect of automated, high-precision output of defect types and spatial locations. Compared to existing rule-driven or template-matching image recognition methods, this approach addresses the issues of poor adaptability, low robustness to environmental changes, and high false positive rates, effectively supporting precise inspections and reliable decision-making by robotic systems in a variety of industrial settings.
[0018] 3. This invention utilizes a collaborative design of wireless communication modules and an identification data packaging mechanism to enable multi-standard wireless real-time upload of defect identification results, including Wi-Fi, Bluetooth, and cellular communication paths. This achieves the technical effect of supporting flexible deployment and remote backhaul in non-fixed network scenarios. Compared to existing quality inspection systems that rely on manual reading, storage, and delayed upload, this solves the pain points of long information feedback chains and untimely task responses, improving the real-time performance and operational efficiency of robotic systems in complex spaces.
[0019] 4. This invention utilizes a power management unit designed for multi-module heterogeneous loads, coupled with dynamic power allocation and safety protection mechanisms, to provide unified power support for the robot body, perception module, computing module, and communication module, achieving the technical effect of continuous and stable power supply and ensuring the coordinated operation of each module. Compared to existing robot platform solutions with fragmented power supply designs and unstable power supply voltages, this solves the technical problems of module susceptibility to failure, discontinuous operation, and frequent inspection interruptions, providing a foundation for efficient inspection processes driven by AI vision. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 This is a diagram of the robot architecture of the AI visual quality inspection system of the present invention; Figure 2 A three-dimensional diagram of the robot body of the present invention; Figure 3 It is a front view of the robot body of the present invention.
[0021] Among them, 1. Robot body; 2. Mobile structure; 21. Mobile wheel chassis; 3. Camera bracket; 31. Rotating bearing base; 32. Multi-degree-of-freedom robotic arm; 4. Industrial camera; 5. Thermal imaging camera. DETAILED DESCRIPTION
[0022] The following is combined with Figure 1 -Attached Figure 3 , the present invention is described in further detail.
[0023] The embodiment of the present invention provides a robot based on an AI visual quality inspection system, The robot body 1 includes a mobile structure 2 and an angle-adjustable camera bracket 3; In this embodiment, the robot body 1 constitutes the core support platform of the AI visual quality inspection system in the present invention, which mainly undertakes functions such as motorized movement, sensing equipment mounting and positioning control.
[0024] A mobile structure 2 is provided at the bottom of the robot body 1, which is used to drive the entire machine to move to a designated position in a factory workshop or inspection area to complete the defect collection and inspection tasks of the target workpiece or structural component.
[0025] The mobile structure 2 is preferably a wheeled chassis 21. This chassis 21 is equipped with four conventional wheels, which offer excellent load capacity and ground adhesion. These wheels are symmetrically arranged in a rectangular pattern at the four corners of the chassis and driven individually by chassis control motors, enabling the robot to perform basic maneuvering functions such as linear movement within a two-dimensional plane, in-place steering, and path following.
[0026] This embodiment utilizes a conventional wheeled chassis 21, which features a simple mechanical structure and high operational stability, making it ideal for industrial quality inspection scenarios requiring path repeatability and stable motion accuracy. The wheels are connected to drive motors for rotational output, which are controlled by a central control unit and work in conjunction with the navigation module to perform path correction and dynamic obstacle avoidance.
[0027] A camera bracket 3 is mounted on top of the robot body 1. This structure is used to secure and adjust the position and angle of the multimodal visual perception device, ensuring clear image information of the target detection area. To improve detection coverage and adapt to complex detection environments, the camera bracket 3 includes a rotating bearing base 31, which is fixedly connected to the top of the robot body 1.
[0028] The swivel bearing base 31 is rotatable, allowing for 360-degree rotation in the horizontal plane, thereby adjusting the camera's orientation. The swivel bearing is driven by a built-in electric rotation unit, and its rotation angle is automatically set by a central control system based on the target position and image recognition requirements.
[0029] A multi-degree-of-freedom robotic arm 32 is mounted on top of the rotating bearing base 31. Preferably, the robotic arm has three or more degrees of freedom, enabling flexible adjustment of the camera module's pitch, telescope, and spatial position. This robotic arm typically utilizes a series or parallel structure, comprised of multiple servo-driven joints. This allows the camera module to be flexibly positioned and adjusted in complex workpiece surfaces or in blind spots, thereby improving image acquisition coverage and accuracy.
[0030] Each joint position of the multi-degree-of-freedom robotic arm 32 is equipped with an angle encoder, and the angle signal is fed back to the robot control system in real time through a feedback loop to form a closed-loop control, thereby ensuring real-time controllable and precise adjustment of the camera posture.
[0031] In one embodiment, the camera bracket 3 is connected to the robot body 1 via a shock-absorbing module, which helps to reduce the impact of vibration on the camera imaging quality during the movement of the robot.
[0032] In this embodiment, the camera bracket 3 can not only rotate at all angles in the horizontal direction, but also achieve multi-angle pitch transformation in the vertical direction under the joint action of the rotating bearing base 31 and the multi-degree-of-freedom mechanical arm 32.
[0033] At least one set of industrial cameras and thermal imaging camera modules is installed on the camera bracket 3, and is used to synchronously collect image data after the robot moves to the target detection area. The image data includes visible light image data and thermal imaging image data. In this embodiment, the industrial camera and thermal imaging camera module constitute the key perception components for image information acquisition in the AI visual quality inspection system of the present invention.
[0034] The at least one camera module is composed of an industrial-grade visible light camera and a thermal imaging camera, which work together to form a dual-modal image acquisition system for acquiring visible light image data and thermal imaging image data in the same detection area.
[0035] To achieve efficient data acquisition and a stable working posture, the industrial camera and thermal imaging camera module are preferably mounted at the end of a multi-degree-of-freedom robotic arm 32 supported by a camera bracket 3. This structural deployment method provides the sensing module with flexible spatial position adjustment capabilities. In inspection tasks involving complex workpiece shapes or limited local space, the robotic arm can actively adjust its position and angle to achieve precise alignment and full coverage of the target area.
[0036] The industrial camera 4 is used to collect visible light image data of the inspected target. It preferably has high resolution, high frame rate and good color reproduction capability, and is suitable for capturing explicit defect features such as target surface texture, contour, scratches, cracks, etc.
[0037] The thermal imaging camera 5 uses infrared band sensing technology to collect thermal radiation image information on the surface of the target object, reflecting the temperature distribution of the measured area, and is used to detect internal structural abnormalities, thermal unevenness or hidden defects.
[0038] To ensure the consistency and integration of collected data, the industrial camera 4 and the thermal imaging camera module achieve data acquisition timing alignment through synchronous triggering control. This synchronization is achieved by the central processing control module issuing a unified signal, controlling the two types of cameras to capture the same target area at the same time, thus ensuring the temporal and spatial correspondence of the collected results.
[0039] In addition, in order to reduce the differences in imaging perspective and focal length between the two types of images, a preset mounting fixture and fixing structure are provided between the camera modules, so that the industrial camera 4 and the thermal imaging camera 5 are fixed at a certain reference position, forming a dual camera path with approximately overlapping fields of view.
[0040] To further ensure the accuracy of subsequent image spatial registration, the two types of cameras are calibrated using a calibration plate or reference pattern during the installation phase to obtain their intrinsic parameters (such as focal length and distortion) and relative extrinsic parameters (such as translation vector and rotation matrix). This calibration process is performed offline as part of system initialization.
[0041] The transformation relationship obtained after calibration can be used for image space correction during actual operation to ensure accurate alignment of images of different modalities at the pixel level, and provide structured aligned input image data for subsequent multimodal feature fusion modules.
[0042] During data collection, the robot 1 moves to the target detection area after receiving the task instruction. Once in position, the robot arm adjusts the spatial orientation of the industrial camera 4 and thermal imaging camera 5 based on the specific posture, size, and detection requirements of the workpiece, ensuring that the sensing field of view covers the entire target area.
[0043] In summary, in this embodiment, at least one set of industrial cameras and thermal imaging camera modules, leveraging the flexible deployment and high-precision synchronization control strategy of the multi-degree-of-freedom robotic arm 32, achieves high-quality, simultaneous acquisition of dual-modal image information. This acquisition method forms the fundamental data source for the present invention's high-precision visual quality inspection.
[0044] The image preprocessing module is connected to the industrial camera and the thermal imaging camera module, and is used to preprocess the collected image data. The preprocessing includes adaptive contrast enhancement, edge sharpening and noise suppression of the collected image data. In this embodiment, the image preprocessing module is arranged after the industrial camera and the thermal imaging camera module. As the front-end basic link in the image processing process of the present invention, its main function is to optimize the structure and improve the quality of the collected original image data, ensuring that the data input into the subsequent image fusion and defect recognition module has good processability and stability.
[0045] The image preprocessing module includes the following functional submodules, each of which is connected through an image data stream to form a processing chain. The processing order is: adaptive contrast enhancement submodule, edge sharpening submodule and noise suppression submodule.
[0046] Adaptive contrast enhancement submodule: In this embodiment, the adaptive contrast enhancement submodule is used to dynamically adjust the grayscale or brightness value of the image to enhance the brightness difference between the target area and the background, thereby improving the overall separability of the image.
[0047] This submodule uses algorithms such as Adaptive Histogram Equalization (AHE) or Contrast Limited Adaptive Histogram Equalization (CLAHE) to achieve brightness enhancement according to the image input type (visible light or thermal imaging image).
[0048] The basic implementation principle is: First, the image is divided into several local areas; Calculate the histogram distribution of each local area separately to form a local contrast model; Remap pixel values according to the histogram mapping function to expand the grayscale dynamic range; If the CLAHE algorithm is used, the contrast enhancement amplitude in the local histogram is also limited to a set threshold to avoid local overexposure or loss of details.
[0049] Let the original image be I(x,y) and the enhanced image be I ′ (x,y), then the enhancement process can be expressed as: Among them, τ AHE / CLAHE is the region adaptive histogram equalization mapping function.
[0050] Through this processing, the brightness differences in the detail areas of the image are amplified, especially for the target areas with defects, which can enhance their grayscale performance and distinguish them from the background.
[0051] Edge sharpening submodule: The edge sharpening submodule is connected to the adaptive contrast enhancement submodule. Its main function is to further enhance the edges and contour structures of objects in the image, especially to improve the recognizability of defect boundaries.
[0052] This module enhances edge information by detecting locations in the image where grayscale values change dramatically. It often uses the Laplacian operator, Sobel gradient operator, or Unsharp Masking algorithm to perform weighted superposition on edge response areas.
[0053] The mathematical expression of the commonly used sharpening operation is: in, represents the second-order derivative operation of the image; λ is the edge enhancement weight parameter; I″(x,y) is the image after edge enhancement.
[0054] In practical applications, this module detects areas with sudden grayscale changes in the image and enhances the contrast of these boundaries, thereby making the edges of the defect morphology and surrounding structures prominent, which helps subsequent image segmentation or recognition algorithms to accurately extract the defect area.
[0055] Noise suppression submodule: In this embodiment, the noise suppression submodule is provided after the edge sharpening submodule to further remove non-structural interference and high-frequency noise signals in the image data, ensuring that the image structural features are accurately preserved without distortion.
[0056] This module mainly solves the problems of image noise and background disturbance caused by sensor thermal noise, ambient light fluctuations or signal transmission errors during image acquisition.
[0057] Preferably, the submodule adopts bilateral filtering or non-local means filtering (NLM) algorithm. Bilateral filtering not only considers the distance between pixels in the spatial neighborhood, but also integrates the similarity of pixel grayscale to achieve simultaneous smoothing of the image and edge preservation.
[0058] The specific calculation formula is as follows: Among them, I ″′ (x, y) represents the pixel value of the output image after bilateral filtering; I ″(i, j) represents the grayscale value of the pixel in the neighborhood of the input image; Ω represents the neighborhood window centered on the pixel (x, y); Represents the spatial distance weight function, which is used to retain the influence of neighboring pixels; Represents the pixel grayscale similarity weight, which is used to retain edge information; W = ∑ (i,j)∈Ω ω s ·ω r Represents the normalization coefficient to ensure that the pixel value after filtering is within a reasonable range.
[0059] The above processing significantly improves the image smoothness.
[0060] In summary, the image preprocessing module in this embodiment adopts a three-level cascade processing structure of "adaptive enhancement-edge enhancement-noise suppression" to optimize and adjust the brightness, edge and noise characteristics of the image in turn to form clear, contrast-rich and interference-free standardized image data.
[0061] The multimodal fusion module is connected to the image preprocessing module and is used to perform spatial registration and feature fusion on the preprocessed image data to generate fused image data; In this embodiment, the multimodal fusion module is arranged after the image preprocessing module. Its main function is to perform multi-step collaborative processing on the visible light image data and thermal imaging image data that have completed preprocessing, including spatial geometric alignment, feature extraction and multi-dimensional feature fusion, and finally generate unified fused image data to support the input requirements of the subsequent defect recognition module.
[0062] This module contains three functional sub-modules: spatial registration sub-module, feature extraction sub-module and feature fusion sub-module, which are connected in sequence to form a fusion processing link.
[0063] Spatial registration submodule: The spatial registration submodule is used to realize the geometric space alignment between images of different modalities, ensuring a one-to-one mapping relationship between thermal imaging images and visible light images at the pixel level.
[0064] To achieve this goal, the spatial registration submodule performs internal and external parameter calibration on the industrial camera and thermal imaging camera respectively during the system initialization phase, obtains the camera model parameters of each imaging mode, including focal length, principal point position and distortion coefficient, and calculates the relative pose matrix T v→t , which is used to describe the spatial transformation relationship between the two modal images. During operation, the registration module performs coordinate transformation on the image according to the calibration results, and often uses affine transformation or perspective transformation model to achieve pixel-level mapping. The transformation formula is: p t =H·p v ; Among them, p v is the pixel coordinate vector in the visible light image; p t is the mapping coordinate in the corresponding thermal imaging image; H is the homography transformation matrix from the visible light image to the thermal imaging image.
[0065] The above coordinate registration process can be completed by matching feature points between images and estimating the transformation matrix. For example, key points are extracted through local invariant feature extractors such as SIFT or ORB, and abnormal matching pairs are eliminated using algorithms such as RANSAC to improve the registration accuracy and robustness.
[0066] After registration, the thermal image and the visible light image remain spatially consistent, ensuring that subsequent feature point information has positional consistency between the two modal images.
[0067] Feature extraction submodule: In this embodiment, the feature extraction submodule is connected to the spatial registration submodule to extract key visual features that are helpful for defect recognition from the geometrically aligned images.
[0068] For the visible light image part, the extracted features mainly include edge features and texture features, and the detailed structure in the image can be extracted through algorithms such as Sobel edge operator, Gabor filter or local binary pattern (LBP).
[0069] For the thermal imaging image part, the extracted features are mainly the spatial gradient characteristics of the temperature distribution map, the contour characteristics or statistical characteristics of the hot spot area, etc.
[0070] Assume that the feature vector of the processed visible light image is expressed as: F visible =ExtractFeatures(I visible ); The feature vector of the thermal imaging image is expressed as: F thermal =ExtractFeatures(I thermal ); Among them, I visible Represents the pre-processed visible light image data; I thermal is the thermal imaging image data after preprocessing; ExtractFeatures(·) is the feature extraction function; F visible is the feature vector extracted from the visible light image, which may include gradient direction histogram, edge contour, local texture distribution, etc.; F thermal The feature vectors extracted from the thermal imaging images may include thermal characteristic indicators such as thermal gradient, hot spot area distribution, and temperature variance.
[0071] The above-mentioned feature vector is usually in the form of a multi-dimensional feature combination, including but not limited to edge intensity histogram, texture direction distribution, thermal gradient direction field and other composite feature sets.
[0072] Feature fusion submodule: The feature fusion submodule is connected to the feature extraction submodule and is used to fuse the feature vectors extracted from the two modal images to construct a unified fusion feature expression.
[0073] The fusion process is completed by weighted feature function addition. The visible light image feature vector and the thermal image feature vector are assigned weights α and β respectively, satisfying the weight normalization condition: α+β=1; The calculation formula of the fused feature vector is: F fusion =α·F visible +β·F thermal ; Among them, F visible Represents the feature vector of the visible light image; F thermal represents the eigenvector of the thermal imaging image, α and β are the weight coefficients of the visible light mode and the thermal imaging mode, respectively, and satisfy α+β=1; F fusion The fusion process of the fused feature vector can adopt a linear weighted form, and can also be expanded to a deep feature fusion structure. In the embodiment of the present invention, a linear fusion method is preferably adopted to ensure that the fusion calculation process is simple, controllable, and explainable.
[0074] The final fusion feature vector F fusion It will be used as unified input data and transmitted to the subsequent defect identification module for further classification and analysis.
[0075] The multimodal fusion module in this embodiment establishes the spatial and semantic association between thermal imaging information and visible light information through three stages of geometric alignment, feature extraction and fusion processing, forming fused image data with unified structure and complementary features.
[0076] A deep learning processor, connected to the multimodal fusion module, is used to perform defect recognition analysis on the fused image data and output defect recognition results including defect type, location, and recognition confidence; In this embodiment, a deep learning processor is installed after the multimodal fusion module. It is primarily used to intelligently analyze and identify defects in the fused image data, constituting the core decision-making component of the present invention's AI-based visual quality inspection system. The processor uses the fused image data as input and, through multi-level deep feature extraction and learning, accurately classifies and locates defects in the target workpiece.
[0077] The deep learning processor includes an image encoding submodule, a defect recognition submodule, and a confidence output submodule. The submodules are interconnected to form an end-to-end recognition and analysis process.
[0078] Image encoding submodule: In this embodiment, the image encoding submodule is used to receive the fused image data I from the multimodal fusion module. fusion , and extract its spatial features and texture features to form a multi-scale feature representation.
[0079] This submodule is preferably constructed using a convolutional neural network (CNN), whose network structure includes multiple convolutional layers, pooling layers, nonlinear activation functions, and skip connection modules. Through layer-by-layer feature abstraction, the encoder can effectively capture the image's edge structure, texture direction, and contextual spatial relationships.
[0080] Assume the input image is The encoded output is: F enc =Encoder(I fusion ); in, Represents a multi-scale feature map; Encoder(·) represents the image encoding network function; d is the channel dimension, which is used to accommodate multi-level semantic features.
[0081] The feature map extracted by this submodule is used to characterize the structural information and potential defect areas in the fused image, providing high-dimensional semantic input for the subsequent recognition module.
[0082] Defect identification submodule: The defect recognition submodule is connected to the image coding submodule, and its core task is to enc The multi-scale feature map represented by the image is used to classify and locate potential defect areas.
[0083] This module preferably uses a deep neural network model for target detection, such as YOLO, SSD, or FasterR-CNN. The model structure generally includes: Candidate box proposal module; Classification branch (predict defect type); Regression branch (predicting defect location coordinates).
[0084] Model outputs include: Defect type label Among them, P cls The category probability vector output by the classification branch; The defect location box b = (x, y, w, h), where x represents the horizontal coordinate of the upper left corner of the defect region bounding box (relative to the input image coordinate system); y represents the vertical coordinate of the upper left corner of the defect region bounding box; w represents the width of the bounding box, i.e., the horizontal pixel range covered by the defect region; and h represents the height of the bounding box, i.e., the vertical pixel range covered by the defect region. This is output by the regression branch and represents the specific spatial region of the defect in the image.
[0085] By performing region detection and feature activation on the coded feature map, the recognition submodule can locate and classify target defects in complex backgrounds.
[0086] Confidence output submodule: In this embodiment, the confidence output submodule is used to quantify the credibility of the output result of the defect identification submodule to provide a complete defect identification result, including the judgment credibility of the quantitative expression model.
[0087] This submodule receives the defect classification probability P cls and positioning accuracy score P loc , and use it as a comprehensive criterion to calculate the final recognition confidence value C. The calculation method uses the Sigmoid function for nonlinear normalization processing, and the formula is as follows: Among them, C represents the recognition confidence of the final output; P cls represents the probability of defect classification; P loc represents the positioning accuracy score; w1 and w2 are preset weight coefficients, and satisfy w1+w2=1; e is the base of the natural logarithm. The confidence value C will serve as a signal measure of the system's recognition effectiveness when performing defect recognition tasks, guiding subsequent result screening or multi-round recognition strategy execution.
[0088] In this embodiment, the deep learning processor uses a modular design to organically combine image feature encoding, defect recognition and determination, and confidence assessment, achieving a highly integrated and scalable intelligent recognition solution. The processor processes the fused image data step by step, not only achieving accurate identification of defects but also enhancing system stability and result reliability through a confidence output mechanism.
[0089] The wireless communication module is connected to the deep learning processor and is used to transmit the defect identification results to the background monitoring system in real time via wireless communication; In this embodiment, the wireless communication module is arranged at the output end of the deep learning processor. Its main function is to send the defect identification results to the background monitoring system in real time via wireless communication, so as to realize remote reporting, centralized management and data retention of detection information.
[0090] The wireless communication module is logically connected to the deep learning processor to receive its output defect recognition data, including defect type, spatial coordinates of the defect in the image, and corresponding recognition confidence score and other result parameters. This information constitutes a complete defect recognition result data packet for subsequent transmission.
[0091] To meet the comprehensive requirements of communication flexibility, environmental adaptability, and bandwidth performance in industrial scenarios, in this embodiment, the wireless communication module preferably includes the following subunits: The Wi-Fi communication unit is used to achieve high-speed data interaction with the background monitoring system in an environment with local area network coverage.
[0092] The Wi-Fi communication unit can establish an encrypted connection with the target network node (such as edge gateway, server) by configuring the standard IEEE802.11 protocol stack, ensuring low-latency transmission of recognition results within a local field without the need for a wired connection.
[0093] The Bluetooth communication unit is used to synchronize the recognition results to the portable terminal or local control module in close-range, point-to-point scenarios.
[0094] The Bluetooth communication unit preferably supports the Bluetooth Low Energy (BLE) protocol to reduce system energy consumption. It is suitable for scenarios with limited space and high mobility, such as mobile quality inspection terminals and wearable operating devices.
[0095] The cellular communication unit is used to upload the recognition results remotely through the public network in an environment without a local area network or under remote deployment conditions.
[0096] The cellular communication unit supports mainstream cellular communication protocols, preferably supporting high-speed mobile networks such as 4G and 5G. In conjunction with the SIM authentication mechanism, it can stably upload recognition results to the cloud database or enterprise monitoring platform, supporting cross-site deployment and centralized quality inspection scheduling.
[0097] In this embodiment, the wireless communication module is internally configured with a data packaging and protocol adaptation component, which is used to organize the recognition result data from the deep learning processor into a unified transmission format according to the following structure: D packet ={ID,Type,Position,Confidence,Timestamp}; Among them, ID represents the image number or detection batch identifier corresponding to the current defect; Type is the defect type code; Position = (x, y, w, h) is the spatial position coordinate of the defect; Confidence = C is the recognition confidence score; Timestamp is the timestamp generated by the current defect recognition.
[0098] The above data structure is encapsulated through the communication protocol stack and pushed in real time using TCP / IP, BLEGATT or cellular-specific APN channels depending on the selected communication mode.
[0099] After receiving data packets, the system automatically decodes, records, and triggers alarms based on a unified data interface. The background monitoring system also dynamically manages upload frequency and communication status, ensuring controllability and robustness of the wireless transmission process.
[0100] In this embodiment, by constructing a communication module that supports multiple wireless standards, it is ensured that the recognition results can achieve stable, secure, and low-latency data transmission in different application scenarios, meeting the functional requirements of the intelligent quality inspection system for remote reporting, intelligent scheduling, and centralized decision-making.
[0101] The power management unit is connected to the robot body, industrial camera and thermal imaging camera modules, image preprocessing module, multimodal fusion module, deep learning processor, and wireless communication module to provide stable operating power; In this embodiment, a power management unit (PMU) is installed within the robot system to continuously provide stable, secure, and adaptive power to all functional modules of the AI-based visual quality inspection system. As a core component of energy supply, this unit's functions extend beyond simple power supply to include multiple sub-functions such as power distribution, voltage conversion, current control, power monitoring, and protection mechanisms, ensuring that each module maintains rated electrical parameters during long-term operation.
[0102] The power management unit establishes physical power transmission paths with the following modules through wired connections: robot body, industrial camera, thermal imaging camera module, image preprocessing module, multimodal fusion module, deep learning processor and wireless communication module.
[0103] To accommodate the heterogeneous electrical requirements of these various subsystems, this embodiment incorporates a multi-channel voltage regulation circuit within the power management unit, with each channel independently corresponding to the target load module. Each output is routed through a DC-DC voltage converter or linear voltage regulator module, enabling adaptive matching of voltage and current levels between the input main power supply and each module.
[0104] Assume the main power input is V in , then the power supply output The actual supply voltage corresponding to the i-th module satisfies: Among them, f i (·) is the voltage conversion function; Each channel is also equipped with overvoltage, overcurrent and short circuit protection mechanisms, which are implemented through segmented current detectors and voltage limiter circuits.i , when: in Indicates the safety current threshold corresponding to this channel. This mechanism prevents cascading power failures caused by downstream module failures.
[0105] In this embodiment, considering the presence of camera acquisition modules (industrial cameras and thermal imaging cameras) and high-performance computing modules (deep learning processors) in the system, whose power consumption characteristics fluctuate greatly, the power management unit integrates a dynamic power allocation algorithm based on load perception, which is used to adjust the output load of each channel according to the real-time operating status of the system.
[0106] The deployment strategy is based on the power demand prediction model P i (t), the power management controller periodically calculates the power consumption: P i (t) = α i ·D i (t)+β i ; Among them, P i (t) is the power consumption prediction value of the i-th module at time t; D i (t) represents the data processing volume or task intensity of the module in the current cycle; α i ,β i is the adjustment parameter obtained by regressing historical power characteristics.
[0107] Through the linkage control of the power prediction mechanism and the output power regulator, the power output has a higher energy efficiency matching, avoiding power waste or insufficient power supply.
[0108] For deep learning processor modules, which are more sensitive to power supply fluctuations, the power management unit preferably integrates a voltage adjustment unit with high voltage regulation accuracy and fast response speed to ensure that the power supply remains stable during short-term load surges (such as model forward reasoning or large-scale image processing).
[0109] In addition, to support the deployment requirements of the system in complex industrial environments, the power management unit is equipped with temperature monitoring and power statistics functions, which can monitor the temperature rise of the power module and the overall power usage in real time. Its monitoring data is transmitted through digital buses (such as I 2 C or SPI) to the control mainboard, which can be used for power consumption analysis and maintenance scheduling of the background system.
[0110] In this embodiment, the power management unit is preferably a highly integrated module in structure, which is easy to embed into the robot body structure, and its input can come from the robot power base or the external industrial power interface.
[0111] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. The robot based on AI visual quality inspection system is characterized by: include: A robot body, the robot body including a camera bracket having a movable structure and an adjustable angle; At least one set of industrial cameras and thermal imaging camera modules, mounted on the camera bracket, for synchronously collecting image data after the robot moves to the target detection area, the image data including visible light image data and thermal imaging image data; An image preprocessing module, connected to the industrial camera and the thermal imaging camera module, for preprocessing the collected image data, wherein the preprocessing includes adaptive contrast enhancement, edge sharpening and noise suppression on the collected image data; A multimodal fusion module, connected to the image preprocessing module, for performing spatial registration and feature fusion on the preprocessed image data to generate fused image data; a deep learning processor, connected to the multimodal fusion module, for performing defect recognition analysis on the fused image data and outputting a defect recognition result including defect type, location, and recognition confidence; A wireless communication module, connected to the deep learning processor, is used to transmit the defect recognition results to the background monitoring system in real time via wireless mode; A power management unit is connected to the robot body, the industrial camera and thermal imaging camera module, the image preprocessing module, the multimodal fusion module, the deep learning processor and the wireless communication module to provide stable working power.
2. The robot based on the AI visual quality inspection system according to claim 1, characterized in that: The mobile structure includes a mobile wheel chassis, which is installed at the bottom of the robot body. The camera bracket includes a rotating bearing base, which is installed at the top of the robot body. A multi-degree-of-freedom robotic arm is installed on the rotating bearing base.
3. The robot based on the AI visual quality inspection system according to claim 1, characterized in that: The industrial camera and thermal imaging camera module are installed on the top of the multi-degree-of-freedom robotic arm.
4. The robot based on the AI visual quality inspection system according to claim 1, characterized in that: The image preprocessing module includes: The adaptive contrast enhancement submodule is used to dynamically adjust the contrast of the collected image data so that the target area and the background area in the image data are clearly distinguished in terms of brightness level; The edge sharpening submodule is connected to the adaptive contrast enhancement submodule to further enhance the edge and contour features of objects in the image to highlight the boundaries of defects; The noise suppression submodule is connected to the edge sharpening submodule and is used to filter out random noise and unstructured interference signals in the image data.
5. The robot based on the AI visual quality inspection system according to claim 1, characterized in that: The multimodal fusion module includes: The spatial registration submodule is used to geometrically align the visible light image and the thermal imaging image in the preprocessed image data to ensure that the visible light image and the thermal imaging image correspond at the pixel level; The feature extraction submodule is connected to the spatial registration submodule and is used to extract key feature information of edges, textures and thermal distribution from the aligned visible light image and thermal imaging image respectively; The feature fusion submodule is connected to the feature extraction submodule and is used to fuse the key features of the visible light image and the thermal imaging image into multi-dimensional information to generate fused image data.
6. The robot based on the AI visual quality inspection system according to claim 5, characterized in that: The feature fusion submodule performs multi-dimensional information fusion on the key features of the visible light image and the thermal imaging image using the following formula: F fusion =α·F visible +β·F thermal ; Among them, F visible Represents the feature vector of the visible light image; F thermal represents the eigenvector of the thermal imaging image, α and β are the weight coefficients of the visible light mode and the thermal imaging mode, respectively, and satisfy α+β=1; F fusion is the fused feature vector.
7. The robot based on the AI visual quality inspection system according to claim 1, characterized in that: The deep learning processor includes: The image encoding submodule is used to receive the fused image data and extract its spatial and texture features to generate multi-scale feature representation; A defect recognition submodule, connected to the image encoding submodule, is used to identify defects in the image data based on the multi-scale feature representation through a pre-trained deep neural network model, and output the defect type and corresponding location coordinates; The confidence output submodule is connected to the defect recognition submodule and is used to extract and output the recognition confidence to form a complete defect recognition result.
8. The robot based on the AI visual quality inspection system according to claim 7, characterized in that: The confidence output submodule calculates the recognition confidence using the following formula: Among them, C represents the recognition confidence of the final output; P cls represents the probability of defect classification; P loc represents the positioning accuracy score; w1 and w2 are preset weight coefficients, and satisfy w1+w2=1; e is the base of the natural logarithm.
9. The robot based on the AI visual quality inspection system according to claim 1, characterized in that: The wireless communication module includes a Wi-Fi communication unit, a Bluetooth communication unit and a cellular communication unit; the wireless communication module transmits the defect identification results to the background monitoring system in real time via wireless means.
Citation Information
Cited By
Honeycomb structure part core lattice defect detection method and system based on robot vision
CN121505560A