Detection system based on body quality inspection robot

By using a embodied quality inspection robot system, which utilizes remote control models and quality inspection models, efficient grasping and inspection of products in different postures can be achieved, solving the problems of inspection accuracy and efficiency in flexible manufacturing, and making it suitable for multi-variety, small-batch production.

CN121010643APending Publication Date: 2025-11-25ZHONGKE HUIYUAN VISUAL TECHNOLOGY (LUOYANG) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510929586.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing product quality inspection equipment is unable to perform efficient quality inspection of products in different postures, especially in flexible manufacturing, where inconsistent product postures lead to decreased inspection accuracy and low efficiency.

Method used

An inspection system based on a body-on-body quality inspection robot is adopted. The main control unit acquires remote operation training datasets and inspection training datasets to train remote operation models and quality inspection models. The robot body performs grasping and inspection based on environmental images, realizing efficient grasping and quality inspection of products in different postures.

Benefits of technology

It improves the efficiency of inspection of products with different postures, can accurately identify structural defects and logical errors, and adapts to the flexible manufacturing needs of multi-variety, small-batch production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121010643A_ABST
    Figure CN121010643A_ABST
Patent Text Reader

Abstract

The invention discloses a detection system based on a quality inspection robot with a body. The detection system comprises a main control unit and at least one robot body, the main control unit obtains a teleoperation training data set and a detection training data set corresponding to the to-be-detected product in different poses, simulative learning training is carried out based on the teleoperation training data set to obtain a teleoperation model, and model training is carried out based on the detection training data set to obtain a quality detection model; the main control unit transmits the remote operation model and the quality detection model to each robot body; and each robot body grabs the target product to at least one target point location based on the collected actual environment image around the target product and the remote operation model, and obtains a quality detection result of the target product based on the collected actual detection image of the target product and the quality detection model. The robot body simulates the grabbing action of a person through self cognition, target products in different poses are grabbed, quality detection is carried out based on the quality detection model, and the detection efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent product testing technology, and in particular to a testing system based on a embodied quality inspection robot. Background Technology

[0002] Product quality inspection is a crucial part of industrial production. With the development of technology, the use of machine vision inspection equipment for product quality inspection has become an important means of industrial production.

[0003] Currently, commonly used machine vision inspection equipment includes a robotic arm, an image acquisition module, and a processing module. Products to be inspected are placed on the inspection table in the same posture. The robotic arm picks up the products in the same posture and moves them to a designated point. The image acquisition module acquires images of the products to be inspected. The processing module extracts features from the acquired images based on image processing algorithms (such as edge detection, template matching, threshold segmentation, etc.) and determines whether there are defects based on the extracted features. Alternatively, the processing module can input the acquired images into a defect recognition model, and the defect recognition model can output the inspection results of the products.

[0004] This method is highly accurate when inspecting products in the same posture. However, when the products have different postures, the robotic arm cannot accurately grasp them to the designated point, thus failing to capture a complete image and reducing inspection accuracy. For example, in flexible manufacturing, due to low production volumes, there is no equipment to position the products in the same posture. Manually arranging them in the same posture would be time-consuming, severely impacting inspection efficiency. Therefore, a product quality inspection device is needed that can efficiently inspect products in different postures. Summary of the Invention

[0005] In view of this, the present invention provides an inspection system based on a body-worn quality inspection robot, the main purpose of which is to solve the problem that existing product quality inspection equipment cannot perform high-efficiency quality inspection of products in different postures.

[0006] According to one aspect of this application, a detection system based on an embodied quality inspection robot is provided, comprising: a main control unit and at least one robot body, wherein the main control unit is electrically connected to each robot body respectively;

[0007] The main control unit acquires the remote operation training dataset and the detection training dataset corresponding to the product under inspection in different poses, performs imitation learning training based on the remote operation training dataset to obtain the remote operation model corresponding to the product under inspection, and performs quality inspection model training based on the detection training dataset to obtain the quality inspection model corresponding to the product under inspection.

[0008] The main control unit transmits the remote operation model and quality inspection model corresponding to the product to be inspected to each of the robot bodies;

[0009] Each robot body acquires actual environmental images around the target product, and based on the actual environmental images and the remote control model, grasps the target product to at least one target location, acquires actual detection images of the target product at each target location, performs quality detection based on the actual detection images and the quality detection model, and obtains the quality detection result of the target product.

[0010] Optionally, each of the robot bodies includes: a robot control unit, a robot torso, a robotic arm, a neck structure, and a head vision component. The robot control unit is connected to the main control unit, the robot torso, the neck structure, the robotic arm, and the head vision component, respectively.

[0011] The robot control unit stores the remote control model and the quality detection model;

[0012] The head vision component is used to acquire environmental images and detect images;

[0013] The neck mechanism is located between the head vision component and the robot torso, and the neck mechanism is used to adjust the pitch angle and left and right head tilting freedom of the head vision component.

[0014] The robotic arm is used to grasp the target product, and the robotic arm is equipped with an image acquisition module for acquiring environmental images.

[0015] Optionally, the inspection system based on the embodied quality inspection robot further includes an isomorphic arm, which is electrically connected to the robotic arm of each robot body. The main control unit acquires a remote operation training dataset and an inspection training dataset corresponding to the product under inspection in different poses, including:

[0016] The main control unit acquires joint motion training data and environmental training images generated when any robot body grasps the product to be inspected in different poses. The joint motion training data is the joint motion data generated by the robotic arm, neck structure and torso of any humanoid robot when the operator wears the isomorphic arm to grasp the product to be inspected in different poses and the robotic arm of any robot body moves synchronously with the isomorphic arm. The environmental training images are environmental images acquired by the image acquisition module and the head vision component.

[0017] For each pose, the main control unit aligns the joint motion training data and environmental training image corresponding to the pose in chronological order, takes the joint motion training data and environmental training image at the same time as a data, sorts the multiple data corresponding to the pose in chronological order, obtains a set of teleoperation training data corresponding to the pose, and uses the multiple sets of teleoperation training data corresponding to the pose as the teleoperation training dataset corresponding to the pose.

[0018] For each pose, the main control unit uses multiple images of the product to be inspected acquired by the vision acquisition component after the product to be inspected is captured to the target position as the detection training dataset.

[0019] Optionally, each of the robot bodies acquires actual environmental images around the target product, and based on the actual environmental images and the remote control model, grasps the target product to at least one target location, and acquires an actual detection image of the target product at the target location, including:

[0020] The image acquisition module and the head vision component acquire the actual environmental image around the target product at the current moment, and input the actual environmental image into the remote control model. The remote control model outputs the joint motion prediction data of the robotic arm, the robot torso and the neck structure at the next moment.

[0021] The robotic arm, the robot torso, and the neck structure perform corresponding actions based on their respective joint motion prediction data for the next moment.

[0022] When the robotic arm picks up the target product and places it at the target location, the head vision component acquires the actual detection image of the target product.

[0023] Optionally, the quality inspection model includes a structural defect detection model and a logic error detection model, and the quality inspection result includes structural defect detection results and logic error detection results; the step of performing quality inspection based on the actual inspection image and the quality inspection model to obtain the quality inspection result of the target product includes:

[0024] The actual detected image is input into the structural defect detection model to obtain the structural defect detection result;

[0025] The actual detected image is input into the logic error detection model to obtain the logic error detection result.

[0026] Optionally, the structural defect detection model can be obtained using the following method:

[0027] Acquire data from multiple modalities related to structural defect detection, and pre-train a pre-defined first visual language large model based on the data from multiple modalities related to structural defect detection to obtain a pre-trained first visual language large model;

[0028] The pre-trained first visual language large model is distilled to obtain the distilled model.

[0029] Multiple detection training images of the product to be inspected are acquired, and the multiple detection training images are input into the distilled model to obtain the first initial detection result.

[0030] The erroneous and undetected results in the first initial detection results are labeled, and the labeled detection results and their corresponding detection training images are used as the first preferred training data.

[0031] Data from multiple modalities related to the first preferred training data are acquired, and the distilled model is trained based on the first preferred training data and the data from the multiple modalities related to it to obtain a structural defect detection model.

[0032] Optionally, the logic error detection model can be obtained using the following method:

[0033] Acquire data from multiple modalities related to logic error detection, and fine-tune the pre-set second visual language big model based on the data from multiple modalities related to logic error detection;

[0034] Acquire multiple detection training images of the product to be inspected, and input the multiple detection training images into the fine-tuned large model to obtain the second initial detection result;

[0035] The erroneous and undetected results in the second initial detection are labeled, and the labeled detection results and their corresponding detection training images are used as the second preferred training data.

[0036] Acquire data from multiple modalities related to the second preferred training data, and train the fine-tuned large model based on the second preferred training data and the data from the multiple modalities related to it to obtain a logic error detection model.

[0037] Optionally, the quality inspection results include qualified, unqualified, and secondary inspection; when the quality inspection results of the target product include secondary inspection, the robot control unit obtains the point corresponding to the secondary inspection from the quality inspection results, and controls the robotic arm to grasp the target product to the point corresponding to the secondary inspection based on the actual environmental image acquired by the image acquisition module and the machine vision component and the remote control model.

[0038] Optionally, the head vision component includes a first camera and a second camera, wherein the detection accuracy of the second camera is greater than that of the first camera; after the robotic arm grasps the target product to the point corresponding to the secondary detection, the head vision component performs image acquisition on the target product according to the secondary detection information, and transmits the image acquired by the secondary detection to the quality detection model.

[0039] Optionally, when there are multiple categories of products to be inspected, the inspection system based on the embodied quality inspection robot further includes:

[0040] The main control unit obtains the remote operation model and quality inspection model corresponding to each of the products to be inspected, and transmits the remote operation model and quality inspection model corresponding to each of the products to be inspected to each of the robot bodies.

[0041] The robot body determines the category of the target product to be inspected through human-computer interaction, and determines the corresponding remote operation model and quality inspection model based on the category of the target product.

[0042] By employing the above-described technical solutions, the technical solutions provided by the embodiments of the present invention have at least the following advantages:

[0043] This application provides an inspection system based on an embodied quality inspection robot, comprising a main control unit and at least one robot body. The main control unit acquires a remote control training dataset and a detection training dataset corresponding to the product to be inspected in different poses. Based on the remote control training dataset, it performs imitation learning training to obtain a remote control model corresponding to the product to be inspected. Based on the detection training dataset, it trains a quality inspection model to obtain a quality inspection model corresponding to the product to be inspected. The main control unit transmits the remote control model and the quality inspection model corresponding to the product to be inspected to each robot body. Each robot body acquires actual environmental images around the target product and inputs these images into the remote control model. The remote control model outputs joint motion data of the robot body. Based on the joint motion data, the robot body grasps the target product to at least one target point and acquires actual inspection images of the target product at each target point. These actual inspection images are then input into the quality inspection model to obtain the quality inspection result of the target product. The robot body, through "embodied cognition," can imitate human grasping actions to grasp the target product in different poses. Based on the quality inspection model and the acquired actual inspection images of the target product, quality inspection is performed, improving inspection efficiency.

[0044] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0045] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0046] Figure 1 A flowchart of a detection system based on a body-worn quality inspection robot provided in an embodiment of this application is shown;

[0047] Figure 2 A schematic diagram of the structure of an inspection system based on a embodied quality inspection robot provided in an embodiment of this application is shown;

[0048] Figure 3 This paper shows another flowchart of a detection system based on a body-mounted quality inspection robot provided in an embodiment of this application;

[0049] Figure 4 The diagram illustrates the structure of an inspection system based on a embodied quality inspection robot, provided in an embodiment of this application, for different customers.

[0050] in,

[0051] In the diagram: 1-Head vision component; 2-Neck mechanism; 3-Robotic arm; 4-Waist mechanism; 5-Automatic navigation system; 6-RGBD camera;

[0052] 10-Data platform; 20-Main control unit; 30-Robot body. Detailed Implementation

[0053] The present invention will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present invention can be combined with each other.

[0054] To further illustrate the technical means and effects adopted by the present invention to achieve the intended purpose, the specific embodiments, structures, features, and effects according to the present invention will be described in detail below with reference to the accompanying drawings and preferred embodiments. In the following description, different "an embodiment" or "an embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0055] To address the problem that existing product quality inspection equipment cannot perform high-accuracy and high-efficiency quality inspection of products in different postures, this application provides an inspection system based on a unibody quality inspection robot, including: a main control unit and at least one robot body, with the main control unit electrically connected to each robot body; the flowchart of the inspection system based on the unibody quality inspection robot is as follows. Figure 1 As shown:

[0056] 102: The main control unit acquires the remote operation training dataset and the detection training dataset corresponding to the product under inspection in different poses, performs imitation learning training based on the remote operation training dataset to obtain the remote operation model corresponding to the product under inspection, and trains the quality inspection model based on the detection training dataset to obtain the quality inspection model corresponding to the product under inspection.

[0057] 104: The main control unit transmits the remote control model and quality inspection model corresponding to the product to be inspected to each robot body;

[0058] 106: Each robot body collects actual environmental images around the target product, grasps the target product to at least one target point based on the actual environmental images and the remote control model, collects actual inspection images of the target product at each target point, performs quality inspection based on the actual inspection images and the quality inspection model, and obtains the quality inspection results of the target product.

[0059] Specifically, rigidly manufactured products are mass-produced products, and are sequentially and uniformly arranged into a fixed position for quality inspection by testing equipment. Flexible manufactured products, on the other hand, are multi-variety, small-batch products. Due to economic reasons, flexible manufactured products are usually not equipped with testing equipment that arranges them into a uniform fixed position. Some flexible manufactured products are difficult to place on the testing table in the same position, and the test images collected during testing are not complete. Therefore, the existing intelligent testing equipment has relatively low accuracy in detecting products with different positions.

[0060] This application proposes a robotic inspection system for quality control, comprising multiple robot bodies and a main control unit. An operator wears a homogeneous arm that moves synchronously with the robotic arm of the robot body. When the operator grasps products in different poses and places them at different target locations, the robotic arm performs synchronized movements, acquiring motion data of each joint of the robot body and simultaneously collecting environmental images around the product. After a product is grasped at each target location, an inspection image of the product is acquired. The joint motion data of the robot body and the acquired environmental images are used as a remote control training dataset, and the inspection image corresponding to each target location is used as an inspection training dataset. Imitation learning is performed based on the remote control training dataset to obtain a trained remote control model. A quality inspection model is then trained based on the inspection training dataset to obtain a trained quality inspection model.

[0061] When inspecting a target product, environmental images surrounding the product are collected. The trained remote-controlled model outputs joint motion data of the robot body based on these images. The robot then executes corresponding actions based on this data, grasping the target product at each target location. Actual inspection images of the target product are collected and input into a quality assessment model to obtain the product's quality inspection results. This allows the robot to grasp target products in different poses by mimicking human grasping actions through "embodied cognition."

[0062] During imitation learning training, frequency domain features of joint motion data are extracted using methods such as Fourier transform and wavelet transform. A sliding window can also be constructed to calculate statistical features such as mean, variance, and peak value within the window. Feature extraction is performed on the environmental image to obtain high-level semantic features, such as the shape, position, and texture of the target product. A suitable imitation learning model architecture is selected, such as behavior cloning, which can directly learn the mapping relationship between the operator's actions and the environmental state.

[0063] Environmental image features and joint motion data features are used as inputs to the behavior cloning model. Optimization algorithms (such as stochastic gradient descent and Adam algorithm) are employed to minimize the error between the model output and the operator's actual movements, continuously adjusting the model parameters. During training, the model is evaluated using a validation set, and hyperparameters, such as the learning rate and the number of network layers, are adjusted based on the evaluation results to prevent overfitting, improve the model's generalization ability, and obtain the final trained telemanipulation model.

[0064] This application provides an inspection system based on an embodied quality inspection robot. Compared with existing technologies, it includes a main control unit and at least one robot body. The main control unit acquires remote operation training datasets and detection training datasets corresponding to the product to be inspected in different poses. Based on the remote operation training datasets, it performs imitation learning training to obtain a remote operation model corresponding to the product to be inspected. Based on the detection training datasets, it trains a quality inspection model to obtain a quality inspection model corresponding to the product to be inspected. The main control unit transmits the remote operation model and quality inspection model corresponding to the product to be inspected to each robot body. Each robot body collects actual environmental images around the target product and inputs these images into the remote operation model. The remote operation model outputs joint motion data of the robot body. Based on the joint motion data, the robot body grasps the target product to at least one target point and collects actual inspection images of the target product at each target point. These actual inspection images are then input into the quality inspection model to obtain the quality inspection result of the target product. The robot body, through "embodied cognition," can imitate human grasping actions to grasp the target product in different poses. Based on the quality inspection model and the collected actual inspection images of the target product, quality inspection is performed, improving inspection efficiency.

[0065] In another embodiment of the invention, each robot body includes: a robot control unit, a robot torso, a robotic arm, a neck structure, and a head vision component. The robot control unit is connected to the main control unit, the robot torso, the neck structure, the robotic arm, and the head vision component, respectively.

[0066] The robot control unit stores remote control models and quality inspection models;

[0067] The head vision component is used to acquire environmental images and detect images;

[0068] The neck mechanism is located between the head vision component and the robot's torso. The neck mechanism is used to adjust the pitch angle and left and right lateral head swing of the head vision component.

[0069] The robotic arm is used to grasp target products. The robotic arm is equipped with an image acquisition module, which is used to acquire environmental images.

[0070] In this embodiment, the robot body consists of a head vision component, a neck mechanism, a robotic arm, and a robot torso. When motion detection is required, the robot body also includes an automatic navigation system, such as... Figure 2 As shown, the robot body consists of a head vision component 1, a neck mechanism 2, a robotic arm 3, a waist mechanism 4, an automatic navigation system 5, and an RGBD camera 6.

[0071] The robot control unit is mounted on the robot's torso. The neck mechanism, located between the head vision unit and the robot's torso, allows adjustment of the head vision unit's pitch and tilt degrees of freedom. The robotic arm, consisting of two arms, connects to the robot's torso and is used to grip the product being tested. The waist structure, located between the robot's torso and the automated guidance system, allows adjustment of the bending degree of freedom in the torso and above. The automated guidance system, located below the robot's torso and in contact with a smooth surface, guides the robot to move in front of the area to be inspected; this automated navigation system can be an AGV (Automated Guided Vehicle).

[0072] In one embodiment of the invention, the inspection system based on the embodied quality inspection robot further includes an isomorphic arm, which is electrically connected to the robotic arm of each robot body. The main control unit acquires remote operation training datasets and inspection training datasets corresponding to the product to be inspected in different poses, including:

[0073] The main control unit acquires joint motion training data and environmental training images generated when any robot body grasps the product under inspection in different poses. The joint motion training data is the joint motion data generated by the robotic arm, neck structure and robot torso of any humanoid robot when the operator wears the isomorphic arm to grasp the product under inspection in different poses and the robotic arm of any robot body moves synchronously with the isomorphic arm. The environmental training images are the environmental images acquired by the image acquisition module and the head vision component.

[0074] For each pose, the main control unit aligns the joint motion training data and environmental training images corresponding to the pose in chronological order, takes the joint motion training data and environmental training images at the same time as a single data set, sorts the multiple data sets corresponding to the pose in chronological order, obtains a set of teleoperation training data corresponding to the pose, and uses the multiple sets of teleoperation training data corresponding to the pose as the teleoperation training dataset corresponding to the pose.

[0075] For each pose, the main control unit uses multiple images of the product to be inspected, acquired by the vision acquisition component after the product to be inspected is captured to the target position, as the detection training dataset.

[0076] In this embodiment, the isomorphic arm is electrically connected to the robotic arm of any robot body. The operator, wearing the isomorphic arm, sequentially grasps the product to be inspected in any pose and places it at each target point. While the operator is grasping the product to be inspected, the robotic arm of the robot body moves synchronously with the isomorphic arm, acquiring joint motion training data of the robotic arm, neck mechanism, and robot torso of the robot body, as well as environmental training images acquired in real time by the head vision component and the image acquisition module on the robotic arm. The joint motion prediction data and environmental training images at the same time are treated as a single data set. The data at the same pose are sorted according to time sequence to obtain a set of teleoperation training data for a pose. Multiple sets of teleoperation training data for a pose are used as the teleoperation training dataset corresponding to that pose.

[0077] Using the above method, the corresponding teleoperation training dataset for each pose is obtained. The teleoperation training dataset for each pose is then input into the initial teleoperation model for imitation learning training, resulting in the trained teleoperation model.

[0078] For each pose, after the product to be inspected is captured at the target point, multiple images of the product to be inspected collected by the vision acquisition component are used as the detection training dataset. The initial quality detection model is trained based on the detection training dataset to obtain the trained quality detection model.

[0079] In one embodiment, the model is trained using reinforcement learning based on the remote control training dataset to obtain the remote control model.

[0080] In another embodiment of the invention, each robot body acquires actual environmental images around the target product, and based on the actual environmental images and the remote control model, grasps the target product to at least one target location, and acquires actual detection images of the target product at the target location, including:

[0081] The image acquisition module and head vision component acquire images of the actual environment around the target product at the current moment, and input the actual environment images into the telecontrol model. The telecontrol model outputs the joint motion prediction data of the robotic arm, robot torso and neck structure at the next moment.

[0082] The robotic arm, robot torso, and neck structure execute corresponding actions based on their respective joint motion prediction data for the next moment;

[0083] When the robotic arm picks up the target product and places it at the target location, the head vision component acquires the actual inspection image of the target product.

[0084] Specifically, when performing quality inspection on a target product in a certain pose, the robot's image acquisition module and head vision component acquire images of the actual environment surrounding the target product at the current moment. This image is then input into the telecontrol model, which outputs the joint motion data for the robotic arm, neck mechanism, and robot torso for the next moment. The robotic arm, neck mechanism, and robot torso execute corresponding actions based on their respective joint motion data. The image acquisition module and head vision component continue to acquire images of the actual environment surrounding the target product at the current moment, and the robotic arm, neck mechanism, and robot torso continue to execute corresponding actions based on the joint motion data output by the telecontrol model for the next moment, sequentially grasping the target product to each target point. The robot body, through "embodied cognition," mimics human grasping actions, enabling it to grasp target products in different poses.

[0085] When the target product is captured at each target point, the head vision component acquires the actual inspection image of the target product and inputs the actual inspection image into the quality inspection model. The quality inspection model performs quality inspection on the target product based on the actual inspection image and outputs the quality inspection result.

[0086] In another embodiment of the invention, the quality inspection model includes a structural defect detection model and a logic error detection model, and the quality inspection result includes structural defect detection result and logic error detection result; quality inspection is performed based on the actual inspection image and the quality inspection model to obtain the quality inspection result of the target product, including:

[0087] The actual detected image is input into the structural defect detection model to obtain the structural defect detection result;

[0088] The actual detected image is input into the logic error detection model to obtain the logic error detection result.

[0089] In this embodiment of the invention, a logic error detection model is used to identify logical errors in the product's appearance, ensuring that the product's appearance conforms to expected logic, such as incorrect graphic colors, missing elements, or incorrect posture. A structural defect detection model is used to identify morphological and dimensional defects in the product, such as cracks and flaws. By using the logic error detection model and the structural defect detection model for targeted detection, all defects in the target product can be accurately identified, improving the accuracy of quality inspection of the target product.

[0090] In this embodiment of the invention, the structural defect detection model is obtained using the following method:

[0091] Acquire data from multiple modalities related to structural defect detection, and pre-train a pre-defined first visual language large model based on the data from multiple modalities related to structural defect detection to obtain a pre-trained first visual language large model;

[0092] The pre-trained first visual language large model is distilled to obtain the distilled model.

[0093] Acquire multiple detection training images of the product to be inspected, input the multiple detection training images into the distilled model, and obtain the first initial detection result;

[0094] The erroneous and undetected results in the first initial detection results are labeled, and the labeled detection results and their corresponding detection training images are used as the first preferred training data.

[0095] Data from multiple modalities related to the first preferred training data are acquired. The distilled model is then trained based on the first preferred training data and the data from the multiple modalities related to it to obtain a structural defect detection model.

[0096] Specifically, the visual language big model refers to a deep learning model that learns rich visual features and semantic information through pre-training on large-scale visual data such as images and videos, thereby enabling it to effectively process various visual tasks. The first visual language big model possesses some general structural defect detection capabilities.

[0097] Data on multiple modalities related to structural errors is acquired, and a large-scale first visual language model is pre-trained based on this data. This pre-trained model allows the large-scale model to learn more general knowledge and improve its ability to detect structural defects.

[0098] Because large models are quite large and computationally intensive, and while training on massive datasets yields a wealth of knowledge, this knowledge may contain redundant information. Model distillation of the pre-trained large visual language model extracts key knowledge and condenses it into a smaller model. This allows the smaller model to express this knowledge more concisely, improving its computational speed and generalization ability.

[0099] Multiple training images of the product to be inspected are acquired. Based on these images, a small-scale model after distillation is trained to obtain initial detection results. These initial results are then reviewed, with annotations added to identify incorrect and undetected results, serving as the first set of preferred training data. This indicates that the small-scale model's ability to identify these errors needs improvement. The preferred training data and related modal data are then input into the small-scale model after distillation for further training, resulting in a trained structural defect detection model. Training the structural defect detection model using a small number of positive and negative samples improves detection accuracy.

[0100] In this embodiment of the invention, the logic error detection model is obtained using the following method:

[0101] Acquire data from multiple modalities related to logic error detection, and fine-tune the pre-set second visual language big model based on the data from multiple modalities related to logic error detection;

[0102] Acquire multiple detection training images of the product to be inspected, input the multiple detection training images into the fine-tuned large model, and obtain the second initial detection result;

[0103] The erroneous and undetected results in the second initial detection are labeled, and the labeled detection results and their corresponding detection training images are used as the second preferred training data.

[0104] Data from multiple modalities related to the second-optimized training data are acquired. Based on the second-optimized training data and the data from the multiple modalities related to it, a fine-tuned large model is trained to obtain a logic error detection model.

[0105] Specifically, the pre-defined second visual language model possesses some general logical error detection capabilities. Data from multiple modalities related to logical errors is acquired, and the second visual language model is fine-tuned based on this data, allowing the model to learn more general knowledge and improve its logical error detection capabilities.

[0106] Multiple training images of the product to be inspected are acquired. Based on these images, a fine-tuned second visual language model is trained to obtain initial detection results. These initial results are then reviewed and annotated, with incorrectly identified and undetected results noted and used as second-optimized training data. This indicates that the fine-tuned second visual language model needs improvement in recognizing these errors. The second-optimized training data and related multimodal data are input into the fine-tuned second visual language model for further training, resulting in a trained logic error detection model. Training the logic error detection model using a small number of positive and negative samples improves detection accuracy.

[0107] In one embodiment, the quality inspection results include qualified, unqualified, and secondary inspection; when the quality inspection results of the target product include secondary inspection, the robot control unit obtains the point corresponding to the secondary inspection from the quality inspection results, and controls the robotic arm to grab the target product to the point corresponding to the secondary inspection based on the actual environmental image and remote operation model collected by the image acquisition module and machine vision component.

[0108] Specifically, after the robotic arm grasps the target product to each target location, the head vision inspection component acquires the actual inspection image of the target product and inputs it into the quality inspection model. The quality inspection model outputs the quality inspection result for each target location, indicating whether the result is qualified, unqualified, or requires secondary inspection. If the inspection result is qualified, the robotic arm directly sorts the target product into the qualified area; if the inspection result is unqualified, the robotic arm directly sorts the target product into the unqualified area.

[0109] If the detection result is a secondary detection, it may be due to unclear areas, requiring a second, more precise inspection. The position coordinates of the target product during the secondary detection are calculated, i.e., the secondary detection point. Based on the environmental images acquired by the image acquisition module and the head vision component, and the secondary detection point, the remote control model outputs joint motion data, enabling the robotic arm to grasp the target product to the secondary detection point. The head vision component then uses a high-precision camera to take a picture.

[0110] In one embodiment, the head vision component includes a first camera and a second camera, wherein the detection accuracy of the second camera is greater than that of the first camera; after the robotic arm grasps the target product to the point corresponding to the secondary detection, the head vision component performs image acquisition on the target product based on the secondary detection information and transmits the image acquired by the secondary detection to the quality inspection model.

[0111] Specifically, the head vision component includes two or more camera lens modules, namely a first camera and a second camera. The first camera is a wide-field-of-view camera, and the second camera is a high-precision camera. The optimal working distance of the different cameras is the same, but the imaging resolution is different, and the focal length of each lens is different.

[0112] The first camera is used for the initial inspection, and the second camera is only used when a second inspection is required. Therefore, when performing a second inspection, it is necessary to calculate the position coordinates of the target product when the second camera is used for inspection.

[0113] First, a workpiece coordinate system is established using the target surface of the first camera. Then, the XY plane position difference is calculated, that is, the spatial position deviation (Δx, Δx) between the first camera and the second camera in the XY plane is calculated. Based on the coordinates of the target product detected by the first camera, the algorithm calculates and outputs the secondary detection position and the coordinate offset (θx, θy) of the first camera imaging center. When the second camera performs fine inspection, the Z-phase compensation for the working distance difference between the two cameras and the XY-phase compensation for the spatial position difference Δ and dynamic offset θ are performed in the workpiece coordinate system to generate the coordinates of the target product during the final secondary fine inspection.

[0114] The detection field of view of the first camera is x∈(X1,X2), y∈(Y1,Y2), and the working distance is within the lens depth of field range z∈(Z1,Z2). That is, the trapezoidal region centered on the center of the first camera is the detectable area of ​​the large-field-of-view camera lens. Based on the target product coordinates in the large-field-of-view image of the first camera (x, y), the center point offset coordinates are:

[0115] θx = x - (X2 - X1) / 2,

[0116] θy=y-(Y2-Y1) / 2

[0117] Let d be the relative position of the center of the first camera and the center of the second camera, i.e., the distance between the center points of the two cameras, and let δ1 be the pixel resolution of the first camera. Calculate the offsets in the x and y directions respectively:

[0118] Δdx=(θx×δ1+d),

[0119] Δdy=(θy×δ1+d)

[0120] The robot arm moves the target product by (Δdx, Δdy) based on the coordinates (x, y) so that the detection area appears at the center of the second camera's field of view. At this time, (x∈(X3,X4),y∈(Y3,Y4)) should be satisfied.

[0121] The second camera acquires high-precision images of the target product in real time. HR Based on the wide-field-of-view image I simultaneously acquired by the first camera IR and high-resolution image I HR Optimize the working distance of the second camera in the Z direction.

[0122] Large field-of-view image I IR Upsampled k times to a high-resolution image I HR For the same pixel size δ2, I is generated IR-up .in:

[0123] k = δ1 / δ2

[0124] Calculate the sum of squared gradients for each of the two images:

[0125] G HR (x,y)=∑(Gx HR (x,y) 2 +Gy HR (x,y) 2 )

[0126] G IR-up (x,y)=∑(Gx(x,y) 2 +Gy(x,y) 2 )

[0127] Where Gx and Gy represent the gradients in the horizontal and vertical directions, respectively, the difference between the squared gradients is used as the evaluation standard for local contrast difference:

[0128] ΔF(x,y)=G HR (x,y)-G IR-up (x,y)

[0129] Calculate ΔF for each pixel in both images. If ΔF is negative in most areas, it indicates that the second camera needs to be refocused to increase contrast. The focusing direction is determined by the contrast-oriented derivative.

[0130] If ΔF'(x,y)>0, then Δdz>0, and z is adjusted in the positive direction; if ΔF'(x,y)<0, then Δdz<0, and z is adjusted in the negative direction.

[0131] The robot's robotic arm moves the target product in the Z direction, fine-tuning the Z-axis distance in real time until ΔF approaches zero or a given value, at which point the fine-tuning stops, satisfying z∈(Z3,Z4). At this point, based on the fine-tuned Z, the second product is moved by the robotic arm to the position corresponding to the secondary detection point, and the second camera performs actual image detection on the target product.

[0132] The area to be tested can be identified as a trapezoidal region centered on the center of the second camera, which simultaneously satisfies the requirement for accurate imaging by both the first and second cameras.

[0133] The second camera can use a low-resolution image of the focused image as a reference in real time during the focusing process, reducing focusing time and lowering the focusing failure rate.

[0134] In one embodiment, when the product to be inspected exists in multiple categories, the inspection system based on the embodied quality inspection robot further includes:

[0135] The main control unit obtains the remote operation model and quality inspection model corresponding to each product to be inspected, and transmits the remote operation model and quality inspection model corresponding to each product to each robot body.

[0136] The robot body determines the category of the target product to be inspected through human-computer interaction, and determines the corresponding remote operation model and quality inspection model based on the category of the target product.

[0137] Specifically, flexible manufacturing, unlike traditional rigid manufacturing (mass production of a single product), is a new production model that arises to address the demand for large-scale customization. It can quickly respond to changes in market demand, product design updates, and manufacturing process variations, and is suitable for multi-variety, small-batch production. Flexible manufacturing involves a wide variety of products, and existing intelligent inspection equipment is traditionally applicable to the inspection of single products, unable to achieve flexible manufacturing suitable for multiple product types. However, the inspection system based on a embodied quality inspection robot proposed in this application can adapt to the diverse product types in flexible manufacturing.

[0138] For each product to be inspected, the operator wears a homogeneous arm to grasp the product in different poses. The robotic arm of any robot body moves synchronously with the homogeneous arm, acquiring joint motion data generated by the robotic arm, neck structure, and torso of that robot body; environmental images acquired by the image acquisition module and head vision component of that robot body; and detection images of the product after it is grasped to the target location, acquired by the vision acquisition component. The joint motion data and environmental images corresponding to the product to be inspected are used as remote control training data for that product. Based on this remote control training dataset, a remote control model for the product to be inspected is trained, resulting in a remote control model for that product. Based on the detection training dataset composed of the detection images of the product to be inspected, a quality inspection model for the product to be inspected is trained, resulting in a quality inspection model for that product. Using the above method, a remote control model and a quality inspection model corresponding to each product to be inspected are obtained. The main control unit then sends the remote control model and quality inspection model corresponding to each product to each robot body.

[0139] The robot body determines the type of target product to be inspected through human-computer interaction, and then determines the remote control model and quality inspection model to use based on the type of target product. The corresponding remote control model and quality inspection model are retrieved, and the robot body grasps the target product within a certain area based on the remote control model, photographs the target product at the target location, performs quality inspection based on the quality inspection model, and sorts the product to different areas based on the quality inspection results. The specific process is as follows: Figure 3 As shown.

[0140] For example, if the test result is qualified, the product is sorted to the qualified area; if the test result is unqualified, the product is sorted to the unqualified area; when the test result requires a second test, the product is picked up and placed at the second test point for a second test.

[0141] In one embodiment, the operator interacts with the robot body, using voice communication to switch between the remote control model and the quality inspection model of the target product.

[0142] like Figure 4As shown, customer A is on the left and customer B is on the right. Their inspection system using embodied quality inspection robots includes a main control unit 20, several robot bodies 30, and a data platform 10. The data platform can obtain remote operation training datasets and inspection training datasets corresponding to the products to be inspected from the main control unit. Based on the remote operation training datasets, it performs imitation learning training to obtain a remote operation model corresponding to the product to be inspected. Based on the inspection training datasets, it trains a quality inspection model to obtain a quality inspection model corresponding to the product to be inspected. The remote operation model and quality inspection model are then transmitted to the main control unit. The data platform also provides data management functions; if customers have any questions, they can log in to the data platform for consultation and feedback.

[0143] The above embodiments are merely exemplary embodiments of this application and are not intended to limit this application. The scope of protection of this application is defined by the claims. Those skilled in the art can make various modifications or equivalent substitutions to this application within its substance and scope of protection, and such modifications or equivalent substitutions should also be considered to fall within the scope of protection of this application.

Claims

1. A detection system based on a embodied quality inspection robot, characterized in that, include: The system includes a main control unit and at least one robot body, wherein the main control unit is electrically connected to each robot body. The main control unit acquires the remote operation training dataset and the detection training dataset corresponding to the product under inspection in different poses, performs imitation learning training based on the remote operation training dataset to obtain the remote operation model corresponding to the product under inspection, and performs quality inspection model training based on the detection training dataset to obtain the quality inspection model corresponding to the product under inspection. The main control unit transmits the remote operation model and quality inspection model corresponding to the product to be inspected to each of the robot bodies; Each robot body acquires actual environmental images around the target product, and based on the actual environmental images and the remote control model, grasps the target product to at least one target location, acquires actual detection images of the target product at each target location, performs quality detection based on the actual detection images and the quality detection model, and obtains the quality detection result of the target product.

2. The inspection system based on a embodied quality inspection robot as described in claim 1, characterized in that, Each robot body includes: a robot control unit, a robot torso, a robotic arm, a neck structure, and a head vision component. The robot control unit is connected to the main control unit, the robot torso, the neck structure, the robotic arm, and the head vision component, respectively. The robot control unit stores the remote control model and the quality detection model; The head vision component is used to acquire environmental images and detect images; The neck mechanism is located between the head vision component and the robot torso, and the neck mechanism is used to adjust the pitch angle and left and right head tilting freedom of the head vision component. The robotic arm is used to grasp the target product, and the robotic arm is equipped with an image acquisition module for acquiring environmental images.

3. The inspection system based on a embodied quality inspection robot as described in claim 2, characterized in that, The inspection system based on the embodied quality inspection robot also includes an isomorphic arm, which is electrically connected to the robotic arm of each robot body. The main control unit acquires remote operation training datasets and inspection training datasets corresponding to the product to be inspected in different poses, including: The main control unit acquires joint motion training data and environmental training images generated when any robot body grasps the product to be inspected in different poses. The joint motion training data is the joint motion data generated by the robotic arm, neck structure and torso of any humanoid robot when the operator wears the isomorphic arm to grasp the product to be inspected in different poses and the robotic arm of any robot body moves synchronously with the isomorphic arm. The environmental training images are environmental images acquired by the image acquisition module and the head vision component. For each pose, the main control unit aligns the joint motion training data and environmental training image corresponding to the pose in chronological order, takes the joint motion training data and environmental training image at the same time as a data, sorts the multiple data corresponding to the pose in chronological order, obtains a set of teleoperation training data corresponding to the pose, and uses the multiple sets of teleoperation training data corresponding to the pose as the teleoperation training dataset corresponding to the pose. For each pose, the main control unit uses multiple images of the product to be inspected acquired by the vision acquisition component after the product to be inspected is captured to the target position as the detection training dataset.

4. The inspection system based on a embodied quality inspection robot as described in claim 2, characterized in that, Each of the robot bodies acquires actual environmental images around the target product, and based on the actual environmental images and the remote control model, grasps the target product to at least one target location, and acquires actual detection images of the target product at the target location, including: The image acquisition module and the head vision component acquire the actual environmental image around the target product at the current moment, and input the actual environmental image into the remote control model. The remote control model outputs the joint motion prediction data of the robotic arm, the robot torso and the neck structure at the next moment. The robotic arm, the robot torso, and the neck structure perform corresponding actions based on their respective joint motion prediction data for the next moment. When the robotic arm picks up the target product and places it at the target location, the head vision component acquires the actual detection image of the target product.

5. The inspection system based on a embodied quality inspection robot as described in claim 1, characterized in that, The quality inspection model includes a structural defect detection model and a logic error detection model, and the quality inspection results include structural defect detection results and logic error detection results. The quality inspection based on the actual detected image and the quality inspection model, to obtain the quality inspection result of the target product, includes: The actual detected image is input into the structural defect detection model to obtain the structural defect detection result; The actual detected image is input into the logic error detection model to obtain the logic error detection result.

6. The inspection system based on a embodied quality inspection robot as described in claim 5, characterized in that, The structural defect detection model was obtained using the following method: Acquire data from multiple modalities related to structural defect detection, and pre-train a pre-defined first visual language large model based on the data from multiple modalities related to structural defect detection to obtain a pre-trained first visual language large model; The pre-trained first visual language large model is distilled to obtain the distilled model. Multiple detection training images of the product to be inspected are acquired, and the multiple detection training images are input into the distilled model to obtain the first initial detection result. The erroneous and undetected results in the first initial detection results are labeled, and the labeled detection results and their corresponding detection training images are used as the first preferred training data. Data from multiple modalities related to the first preferred training data are acquired, and the distilled model is trained based on the first preferred training data and the data from the multiple modalities related to it to obtain a structural defect detection model.

7. The inspection system based on a embodied quality inspection robot as described in claim 5, characterized in that, The logic error detection model is obtained using the following method: Acquire data from multiple modalities related to logic error detection, and fine-tune the pre-set second visual language big model based on the data from multiple modalities related to logic error detection; Acquire multiple detection training images of the product to be inspected, and input the multiple detection training images into the fine-tuned large model to obtain the second initial detection result; The erroneous and undetected results in the second initial detection are labeled, and the labeled detection results and their corresponding detection training images are used as the second preferred training data. Acquire data from multiple modalities related to the second preferred training data, and train the fine-tuned large model based on the second preferred training data and the data from the multiple modalities related to it to obtain a logic error detection model.

8. The inspection system based on a embodied quality inspection robot as described in claim 5, characterized in that, The quality inspection results include qualified, unqualified, and secondary inspection results. When the quality inspection results of the target product include secondary inspection results, the robot control unit obtains the point corresponding to the secondary inspection from the quality inspection results, and controls the robotic arm to grab the target product to the point corresponding to the secondary inspection based on the actual environment image acquired by the image acquisition module and the machine vision component and the remote control model.

9. The inspection system based on a embodied quality inspection robot as described in claim 8, characterized in that, The head vision component includes a first camera and a second camera, wherein the detection accuracy of the second camera is greater than that of the first camera; after the robotic arm grasps the target product to the point corresponding to the secondary detection, the head vision component performs image acquisition on the target product based on the secondary detection information, and transmits the image acquired by the secondary detection to the quality detection model.

10. The inspection system based on a embodied quality inspection robot as described in any one of claims 1-9, characterized in that, When the products to be inspected fall into multiple categories, the inspection system based on the embodied quality inspection robot further includes: The main control unit obtains the remote operation model and quality inspection model corresponding to each of the products to be inspected, and transmits the remote operation model and quality inspection model corresponding to each of the products to be inspected to each of the robot bodies. The robot body determines the category of the target product to be inspected through human-computer interaction, and determines the corresponding remote operation model and quality inspection model based on the category of the target product.