Method for improving 3D machine vision inspection and measurement precision based on large model
Through the combination of multimodal sensors and adaptive deep learning models, the technical bottlenecks of 3D machine vision technology in complex industrial scenarios are solved, high-precision three-dimensional measurement and inspection are realized, the adaptability and stability of the system are improved, and the high-precision detection needs of industrial production lines are met.
Patent Information
- Application Number
- CN202510015736.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-06
AI Technical Summary
The existing 3D machine vision technology has technical bottlenecks in data acquisition integrity, algorithm adaptability, model migration, real-time and accuracy balance, multimodal data fusion and system dynamic adjustment capabilities in complex industrial scenarios, and it is difficult to meet the high precision and high stability requirements of industrial production lines.
Multimodal sensors are used to combine adaptive deep learning models and transfer learning technology to collect three-dimensional data through multimodal sensors. The adaptive deep learning model adjusts the model structure in real time to adapt to complex environments. Transfer learning technology reduces the need for labeled data and enhances the generalization ability of model to achieve high-precision three-dimensional measurement and inspection.
It significantly improves the inspection accuracy and measurement efficiency of 3D machine vision, enhances the adaptability and stability of the system, and can achieve high-precision inspection in complex industrial scenarios, meeting the high-speed operation needs of industrial production lines.
Smart Images

Figure CN119941673A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine vision, and in particular to a method for improving 3D machine vision inspection and measurement accuracy based on a large model. Background Art
[0002] With the rapid development of industrial automation and intelligent manufacturing, 3D machine vision technology has become an indispensable tool in modern industrial production. In the fields of industrial quality inspection, assembly guidance, production line monitoring, etc., 3D machine vision technology can achieve accurate measurement and inspection of targets by obtaining three-dimensional information of objects. However, the existing 3D machine vision technology still has many limitations and technical bottlenecks in practical applications, which restricts its comprehensive promotion and efficient application in complex industrial scenarios.
[0003] Currently, many 3D machine vision systems rely on a single type of sensor (such as a laser scanner or stereo camera) to collect 3D data. Although this single sensor mode can provide stable data support in specific environments, it is often easily disturbed by factors such as ambient lighting, target surface material or texture loss in a changing industrial environment, thus affecting the integrity and reliability of data collection. For example, for highly reflective or textureless surfaces, traditional stereo vision methods may have difficulty generating accurate 3D data, further limiting the accuracy of the measurement.
[0004] At the data processing level, current 3D data processing methods mostly rely on rule-based algorithms (such as feature matching and point cloud stitching), which are usually based on manually designed feature extraction rules. However, when the target object has complex geometric shapes, smooth surfaces, blurred textures, or is in harsh environmental conditions such as low light and high dynamic range, the adaptability and robustness of these algorithms are often insufficient, and it is difficult to meet the requirements of high precision and high stability in industrial scenarios. At the same time, with the development of deep learning technology, some 3D machine vision systems have begun to use deep learning methods for feature extraction and target detection, but such data-driven methods usually require a large amount of labeled data as support. In industrial production environments, detection tasks are diverse, target objects are of various types, and production environments are complex, making the cost of obtaining labeled data extremely high. In addition, the equipment configuration and environmental conditions of different production lines vary greatly, making it difficult to directly migrate and apply existing training models, increasing the difficulty and cost of model deployment.
[0005] Industrial production lines have placed higher demands on the real-time performance and accuracy of 3D machine vision systems. However, the computational complexity of traditional algorithms and some deep learning models is high. When processing large-scale point cloud data or high-resolution images, the system's processing speed is difficult to meet the high-speed operation requirements of the production line. This performance bottleneck is particularly significant in real-time application environments, further limiting its widespread application in industrial scenarios.
[0006] In addition, the development of multimodal data fusion technology is still immature. Industrial applications usually need to face a variety of complex scenarios, and relying on a single type of three-dimensional data cannot fully reflect the characteristics of the target object. Although multimodal sensor fusion technology has made certain progress in recent years, in industrial practice, the data collection cycles between different sensors are different and the accuracy differences are significant. How to efficiently fuse these heterogeneous data is still a technical difficulty. At the same time, the redundancy and consistency issues of multimodal data also put forward higher requirements for the realization of high-precision detection.
[0007] Traditional 3D machine vision systems are usually designed statically, that is, they run after setting fixed parameters in a specific scene or target range. When the type of target object or production environment changes, the existing system often needs to be reconfigured or retrained, lacking the ability to dynamically adjust. This static characteristic not only limits the versatility of the system, but also poses a severe challenge to its ability to adapt to new environments.
[0008] In complex industrial environments, external interference such as lighting changes, dust, and vibration often affects the performance of 3D machine vision systems. For example, laser scanners may generate noise under strong light interference, resulting in reduced reliability of scan data; stereo cameras have difficulty extracting accurate depth information in insufficient light. These problems seriously restrict the stability of 3D machine vision systems, making them face more uncertainties in actual industrial applications.
[0009] Overall, the existing 3D machine vision technology has technical bottlenecks in data collection integrity, algorithm adaptability, model portability, real-time and accuracy balance, multimodal data fusion and system dynamic adjustment capabilities. New technical solutions are urgently needed to break through these limitations and better meet the complex needs of industrial scenarios. Summary of the invention
[0010] In order to solve the above problems in the prior art, the present invention proposes a method for improving 3D machine vision inspection and measurement accuracy based on a large model, which is characterized by comprising the following steps:
[0011] Step S1: collecting three-dimensional data of a target object by means of a multimodal sensor, wherein the multimodal sensor is used to obtain three-dimensional spatial information of the target object;
[0012] Step S2: Processing the three-dimensional data using an adaptive deep learning model to extract three-dimensional features of the target object, wherein the adaptive deep learning model can adjust model structure parameters in real time according to the three-dimensional data and environmental conditions;
[0013] Step S3: Based on the three-dimensional features extracted by the adaptive deep learning model, the knowledge acquired by the model in other environments is transferred to the current target object detection task through transfer learning technology, and the transferred model is fine-tuned;
[0014] Step S4: performing three-dimensional measurement and inspection on the three-dimensional features of the target object according to the fine-tuned deep learning model to generate position, size and shape data of the target object;
[0015] Step S5: Feedback the three-dimensional measurement and inspection results to the industrial automation system.
[0016] The multimodal sensor includes at least one of a laser scanner, a stereo camera or a depth sensor.
[0017] The laser scanner is used to collect high-precision depth information of the target object, and the stereo camera is used to collect texture and surface details of the target object.
[0018] The adaptive deep learning model includes a dynamic adjustment module, which monitors environmental conditions through an environmental detector and adjusts the convolution kernel size, the number of network layers and the activation function type in real time according to the environmental conditions.
[0019] The environmental conditions monitored by the environmental detector include light intensity, background complexity, and reflectivity of the target object surface.
[0020] The three-dimensional measurement includes calculating the length, width and height of the target object through an outer envelope box, and the outer envelope box is generated by extreme points of the point cloud.
[0021] The three-dimensional inspection includes analyzing the shape characteristics of the target object and extracting the shape parameters of the target geometry through a least squares fitting algorithm.
[0022] The industrial automation system includes a quality detection module, which automatically determines whether the target object meets the preset specifications according to the measurement results, and performs a rejection operation on the target object that does not meet the requirements.
[0023] The rejection operation is performed by a pneumatic device, a robotic arm or a diversion device, and the detection results and rejection information of unqualified objects are recorded for subsequent tracking and statistical analysis.
[0024] Beneficial effects:
[0025] The present invention realizes efficient acquisition, accurate feature extraction and intelligent measurement of three-dimensional data of target objects through the combination of multimodal sensors, adaptive deep learning models and transfer learning technology. The adaptive model dynamically adjusts parameters to improve adaptability to complex environments, and transfer learning reduces the need for labeled data and enhances the generalization ability of the model. Combined with real-time feedback from industrial automation systems, the present invention significantly improves the inspection accuracy and measurement efficiency of 3D machine vision, meeting the high-precision detection requirements in complex industrial scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present application, but do not constitute an improper limitation of the present invention. In the drawings:
[0027] Figure 1 A flow chart of the method of the present invention is shown. DETAILED DESCRIPTION
[0028] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments, wherein the illustrative embodiments and descriptions are only used to explain the present invention but are not intended to limit the present invention.
[0029] The present invention provides a method for improving the accuracy of 3D machine vision inspection and measurement based on a large model. By adopting a multimodal sensor, an adaptive deep learning model and transfer learning technology, high-precision three-dimensional measurement and inspection of a target object is achieved, and the measurement results are used in an industrial automation system. The method of the present invention is described in detail below in conjunction with specific embodiments.
[0030] Step S1: 3D data acquisition
[0031] In this embodiment, in order to obtain high-precision three-dimensional data of the target object, a multimodal sensor system is used to comprehensively collect the target object. The multimodal sensor system includes a laser scanner and a stereo camera, which work synchronously through a preset time synchronization module to ensure the accuracy and consistency of data collection.
[0032] The laser scanner is mounted on a stable bracket and scans the target object at multiple angles. Specifically, the laser scanner uses its built-in laser beam projection system to project laser light onto the surface of the target object, and receives the reflected light signal through the built-in optical sensor to generate depth data of the target object. The depth data is converted into a depth map through a processing unit with an accuracy range of 0.01 mm to 0.1 mm. The frequency and sampling rate of the laser beam can be adjusted according to specific application requirements to meet different accuracy requirements. The working parameters of the laser scanner include:
[0033] Laser beam wavelength: 650 nanometers;
[0034] Scanning angle range: ±60 degrees;
[0035] Sampling frequency: 1000 times / second.
[0036] The control signal of the laser scanner is generated by the data acquisition module and sent to the laser scanner drive unit through the industrial protocol to ensure precise control of the scanning path and frequency.
[0037] The stereo camera is mounted on a bracket parallel to the laser scanner to obtain the texture and surface feature information of the target object. The stereo camera includes two high-resolution imaging units, which capture images of the target object from different angles and generate three-dimensional texture data related to the target object through a stereo matching algorithm. The working parameters of the stereo camera include:
[0038] Image resolution: 1920 × 1080 pixels;
[0039] Frame rate: 60 frames per second;
[0040] Lens focal length: 50 mm.
[0041] The camera's exposure time and gain parameters are automatically adjusted according to the acquisition environment. For example, in a highly reflective environment, the camera reduces the exposure time through the built-in exposure adjustment module to avoid overexposure; in a low-light environment, the gain parameter is automatically increased to enhance texture details.
[0042] In order to ensure the timing consistency of the data collected by the laser scanner and the stereo camera, the system is equipped with a hardware-based time synchronization module. This module uses a high-precision clock chip (nanosecond accuracy) as the main clock source, and drives the synchronous start of the laser scanner and the stereo camera through a synchronous pulse signal (Trigger signal). The synchronous pulse signal is generated by the FPGA module with a period of 1 millisecond, and the precise distribution of the trigger signal is achieved through clock division. The synchronization module includes the following key components:
[0043] Main clock chip: built-in high-precision oscillator, frequency stability is ±5ppm;
[0044] FPGA control unit: generates synchronous trigger signals and controls the data acquisition cycle;
[0045] Synchronous interface: Use standard LVDS signals to connect to laser scanners and stereo cameras.
[0046] After the laser scanner and stereo camera complete the acquisition, the generated depth map and texture map are transmitted to the data processing unit through the high-speed data bus. In the data processing unit, the depth map and texture map are fused using a spatial alignment algorithm to ensure that the corresponding relationship between the two in the three-dimensional space is consistent. The fused three-dimensional data set is output in the form of a point cloud. The point cloud format is the standard PLY (Polygon File Format) or OBJ format, which contains the coordinates (X, Y, Z) and color (R, G, B) information of each point.
[0047] The laser scanner provides high-precision depth data, while the stereo camera supplements the texture information of the target surface to ensure the comprehensiveness of the collected data. Both the laser scanner and the stereo camera have the ability to adapt to the environment and can dynamically adjust the collection parameters according to the actual collection scene. The hardware-based time synchronization module provides nanosecond-level synchronization accuracy to ensure the consistency of multimodal data in the time dimension. The fused 3D data adopts a universal point cloud format to facilitate subsequent processing and analysis.
[0048] Through the above technical means, this embodiment realizes the comprehensive collection of high-precision three-dimensional data of the target object, laying a foundation for subsequent feature extraction, inspection and measurement.
[0049] Step S2: 3D feature extraction
[0050] In this embodiment, the collected three-dimensional data is input into an adaptive deep learning model for processing to extract the three-dimensional features of the target object. The adaptive deep learning model is built based on a convolutional neural network (CNN) architecture and can automatically learn and extract effective features from three-dimensional data, thereby providing basic data support for subsequent inspection and measurement steps.
[0051] The adaptive deep learning model adopts a multi-layer convolutional network structure, including an input layer, several convolutional layers, a pooling layer, and a fully connected layer. The input layer receives preprocessed three-dimensional data (such as depth maps and texture maps) and converts them into a format suitable for model calculation. The convolutional layer extracts the spatial features of the target object, the pooling layer is used for dimensionality reduction to reduce computational complexity, and the fully connected layer maps the extracted features into a fixed-length feature vector for use in subsequent steps.
[0052] Specifically, the key layers of the model are configured as follows:
[0053] The convolution layer uses a 3×3 or 5×5 convolution kernel with a stride of 1 and a ReLU (Rectified Linear Unit) activation function.
[0054] The pooling layer uses a maximum pooling operation of 2×2 size to reduce the size of the feature map while retaining the main features.
[0055] The dimension of the feature vector output by the fully connected layer is set according to the specific application requirements, and is commonly set to 256 or 512 dimensions.
[0056] To improve the model's environmental adaptability, the adaptive deep learning model is equipped with a dynamic adjustment module. The dynamic adjustment module consists of an environmental detector and a parameter adjustment unit. Its specific functions are as follows:
[0057] The environmental detector uses sensors to monitor the environmental conditions during data collection in real time, including parameters such as light intensity, reflectivity of the target object surface, and background complexity. For example, by collecting the illumination distribution histogram of the depth map and the contrast value of the texture map, the indicators of the uniformity of the ambient illumination and the characteristics of the target surface are calculated. The detected environmental information is transmitted to the parameter adjustment unit in the form of numerical values.
[0058] The parameter adjustment unit dynamically adjusts the key parameters of the adaptive deep learning model according to preset rules, mainly including the convolution kernel size, the number of network layers, and the activation function type. The adjustment logic is based on the following rules:
[0059] Low-light conditions: Increase the size of the convolution kernel (such as from 3×3 to 5×5) to capture more details, and increase the number of network layers to improve feature extraction capabilities.
[0060] Highly reflective surface: Adjust the activation function to Sigmoid to enhance the response to subtle features, while reducing the pooling operation to retain more local information.
[0061] Texture-rich environment: Reduce the size of the convolution kernel (such as from 5×5 to 3×3), and reduce the number of network layers to improve processing speed.
[0062] The implementation of parameter adjustment is based on the dynamic parameter control algorithm. The specific formula is as follows:
[0063] K new =K default +ΔK·f(E)
[0064] in:
[0065] K new is the adjusted convolution kernel size;
[0066] K default is the default convolution kernel size;
[0067] ΔK is the adjustment increment, usually ranging from 1 to 2;
[0068] f(E) is the environmental characteristic function, which is calculated based on environmental parameters such as light intensity and surface characteristics and is defined as follows:
[0069]
[0070] Where E light Indicates the light intensity index, ranging from 0 to 1.
[0071] Feature extraction process flow:
[0072] Input 3D data, including depth map and texture map, are passed to the input layer after normalization;
[0073] The environmental detector analyzes the environmental characteristics of the input data in real time and transmits the detection results to the parameter adjustment unit;
[0074] The parameter adjustment unit dynamically adjusts the convolution kernel size, the number of network layers, and the activation function type to generate an optimized model structure;
[0075] The adaptive deep learning model extracts three-dimensional features based on the optimized parameter structure and generates feature vectors;
[0076] The extracted feature vectors are output to the subsequent steps.
[0077] Step S3: Transfer learning and fine-tuning
[0078] In this embodiment, in order to reduce the demand for labeled data and improve the adaptability and performance of the model in the new environment, transfer learning technology is used to transfer the knowledge of the pre-trained deep learning model to the target production environment, and the model is fine-tuned through a small amount of labeled data to adapt it to the specific needs of the current production line and the characteristics of the target object.
[0079] This example uses a pre-trained deep learning model from an industry standard dataset as the base model. This model is built on a convolutional neural network (CNN) architecture and has been fully trained on common 3D object detection tasks to extract high-quality 3D features. The specific configuration of the base model is as follows:
[0080] The input layer supports multi-modal data input of depth map and texture map;
[0081] The feature extraction network contains multiple convolutional layers and pooling layers;
[0082] The output layer contains target classification nodes and position regression nodes.
[0083] The training data of the basic model includes a wide range of industrial objects (such as parts, device surfaces, etc.) and their three-dimensional data, covering common geometric shapes and surface material properties, providing a strong general knowledge foundation for transfer learning.
[0084] Transfer learning process:
[0085] At the beginning of transfer learning, the basic feature extraction layer (low-level convolutional network) of the pre-trained model is frozen and its original weights are retained to maintain the general 3D feature extraction capability. These features are highly universal for most target objects and do not need to be readjusted during the transfer process.
[0086] Add a new task-specific layer on top of the base model to adapt the detection requirements of the target production environment. For example, add a shape parameter prediction node based on the target position regression, or adjust the number and type of classification nodes according to the target characteristics.
[0087] For the new task layer, the initial weights are generated by random initialization, and a small amount of labeled data from the target production environment is used for initial training. This process optimizes the parameters of the new task layer through back propagation and gradually adapts to the new environment.
[0088] After completing the preliminary training of the new task layer, unfreeze the parameters of some basic feature extraction layers (such as the last few layers of convolutional networks) and jointly fine-tune the entire model. During the fine-tuning process, the learning rate is dynamically adjusted according to the characteristics of the target environment. Usually, a smaller learning rate (such as 0.0001) is selected for the basic feature extraction layer, while a larger learning rate (such as 0.001) is used for the new task layer to ensure efficient adaptation of the model in the new task.
[0089] The fine-tuning process uses the Mini-Batch Gradient Descent (MBGD) method to optimize the objective function to minimize the target detection error. The specific objective function L′ is as follows:
[0090] L′=α·L cls +β·L reg
[0091] in:
[0092] L cls For classification loss, the cross entropy loss function is used;
[0093] L reg For position regression loss, the L2 loss function is used;
[0094] α and β are weight coefficients used to balance the impact of classification and regression tasks, and are usually set to 1:1.
[0095] During the fine-tuning process, the system dynamically samples the labeled data of the new environment and combines it with transfer learning technology to improve the model's specific detection capabilities for the target environment. For example, when detecting defects on metal surfaces, the model can more accurately identify dents and scratches through fine-tuning while maintaining universality for complex geometric shapes.
[0096] After fine-tuning, the performance of the fine-tuned model in the target environment is evaluated through cross-validation, including target detection accuracy (Precision), recall rate (Recall) and F1 score. The workpieces in the target production environment are used as test objects to verify its application effect on the actual production line. The test shows that transfer learning and the fine-tuned model can significantly improve the detection accuracy of the target object, while reducing the dependence on large-scale labeled data.
[0097] This embodiment successfully achieves rapid adaptation of the basic model to the target environment through transfer learning and fine-tuning technology, providing an effective solution for high-precision target detection and measurement in complex industrial production environments.
[0098] Step S4: 3D measurement and inspection
[0099] In this embodiment, the deep learning model that has undergone transfer learning and fine-tuning optimization is used to perform three-dimensional measurement and inspection of the target object, aiming to generate the position, size, and shape data of the target object, ensuring comprehensive coverage and accurate extraction of key features of complex-shaped objects.
[0100] 3D data input and preprocessing:
[0101] The collected depth map and texture map are passed as input data to the optimized deep learning model. In order to improve the accuracy of measurement and inspection, the input data is first preprocessed, including the following steps:
[0102] Depth map normalization: Map the depth values in the depth map to the range of [0,1] to eliminate the measurement scale differences between different sensors.
[0103] Texture enhancement: Enhance the contrast of texture images through histogram equalization to highlight the surface features of the target object and reduce data loss under high reflection or low light conditions.
[0104] Coordinate alignment and fusion: Use the spatial alignment algorithm to unify the coordinate systems of the depth map and texture map to ensure data consistency and accuracy. The aligned data is used to generate a 3D point cloud of the target object through a point cloud fusion algorithm.
[0105] 3D feature extraction and key feature identification:
[0106] The optimized deep learning model processes the point cloud data to extract the 3D features of the target object and identify key geometric features. The model's feature extraction network includes multiple layers of convolution operations, which can extract the spatial shape characteristics and local geometric features of the object, such as curvature, edges, and vertices. The key feature identification process is based on the following steps:
[0107] Curvature calculation: The key surface area of the object is identified by the principal curvature calculation formula of the local neighborhood of the point cloud. The specific formula is:
[0108]
[0109] This formula is used to calculate the principal curvature of a local area in a point cloud to identify the key surface area of the target object. The principal curvature is one of the important features that characterize the local geometry of a point cloud and can reflect the degree of curvature of the object surface in a certain direction.
[0110] The formula parameters are described as follows:
[0111] C: represents the curvature value, which is the result of the calculation of this formula. The value range of C is between [0,1] and is used to measure the surface characteristics of the point cloud neighborhood:
[0112] When C is close to 0, it means that the surface of the area is relatively flat;
[0113] When C is close to 1, it means that the region is a sharp feature, such as an edge or a vertex.
[0114] λ max : Indicates the maximum eigenvalue (maximum principal curvature) of the main direction of the local neighborhood of the point cloud.
[0115] It is the largest eigenvalue in the covariance matrix of the neighborhood point cloud and usually represents the intensity of change of the local surface in the direction of maximum curvature.
[0116] λ min : Indicates the minimum eigenvalue (minimum principal curvature) of the main direction of the local neighborhood of the point cloud.
[0117] It is the smallest eigenvalue in the covariance matrix of the neighborhood point cloud and usually reflects the intensity of change of the local surface in the direction of minimum curvature.
[0118] Neighborhood point cloud: refers to a local point cloud set defined by radius r or number of points k with the target point as the center. The main direction eigenvalue λ is determined by calculating the covariance matrix of these neighborhood points. max and λ min .
[0119] This formula quantitatively expresses the curvature of an object's surface through the ratio of the principal curvatures C.
[0120] During point cloud processing, the curvature value C is used to distinguish flat areas, curved areas, and sharp areas (such as edges and vertices) of an object.
[0121] A larger curvature value C indicates that the point is located at a significant feature position on the surface of the object and is suitable for further processing as a key feature point.
[0122] This formula, combined with the calculation of local features of point clouds, can effectively support three-dimensional feature extraction and key point recognition of objects with complex shapes.
[0123] Edge detection: Use a gradient-based edge detection algorithm to identify the boundary area of the target object and extract its contour features.
[0124] Key point extraction: Calculate key points based on geometric discontinuities in the point cloud, such as cusps or intersections, for subsequent measurement and analysis.
[0125] 3D measurement and data output:
[0126] When measuring the position, size and shape of the target object, the system generates high-precision measurement data based on the spatial distribution and geometric characteristics of the point cloud. The specific process is as follows:
[0127] Position measurement: Use the point cloud centroid to calculate the 3D position of the target object. c ,Y c ,Z c ) is calculated as:
[0128]
[0129] Parameter Description:
[0130] X c ,Y c ,Z c : Represent the centroid coordinates of the target object in three-dimensional space. The centroid is the geometric center of the object point cloud and represents the position of the target object.
[0131] N: The total number of points in the point cloud. It is used to calculate the average of all points in the point cloud, ensuring that the centroid position is closely related to the distribution of all points.
[0132] x i ,y i ,z i : The coordinates of the i-th point in the three-dimensional coordinate system. The position of each point in the point cloud is represented by these coordinates.
[0133] This formula generates the centroid coordinates (X c ,Y c ,Z c ). The center of mass reflects the position of the target object in three-dimensional space and is the key to the overall positioning of the target object.
[0134] Size measurement: Calculate the length, width and height of the target object through the bounding box of the point cloud. The bounding box is generated by the extreme points of the point cloud, and the formula is as follows:
[0135] L=X max -X min ,W=Y max -Y min ,H=Z max -Z min
[0136] Parameter Description:
[0137] L, W, H: represent the length, width and height of the target object respectively, which is the three-dimensional representation of the size of the target object.
[0138] X max ,X min : The maximum and minimum values of the point cloud in the X direction correspond to the two farthest points of the point cloud along the X axis. L is calculated by the difference between these two extreme points and represents the size of the target object in the X direction.
[0139] Y max ,Y min : The maximum and minimum values of the point cloud in the Y direction correspond to the two farthest points of the point cloud along the Y axis. W is calculated by the difference between these two extreme points and represents the size of the target object in the Y direction.
[0140] Z max ,Z min : The maximum and minimum values of the point cloud in the Z direction correspond to the two farthest points of the point cloud along the Z axis. H is calculated by the difference between these two extreme points and represents the size of the target object in the Z direction.
[0141] This formula generates the three-dimensional size of the target object by calculating the extreme point difference of the point cloud on the three-dimensional coordinate axis. It is used to quickly approximate the volume and spatial range of the target object.
[0142] Shape analysis: Extract the shape characteristics of the target object, such as a plane, cylinder or sphere, based on the point cloud fitting algorithm. Use the least squares fitting method to generate a mathematical model of the target geometry and calculate its specific parameters (such as radius, tilt angle, etc.).
[0143] Comprehensive analysis and inspection:
[0144] After the measurement is completed, the system compares and analyzes the actual measurement data of the target object with the preset standard to identify potential deviations or defects. For example:
[0145] Check whether the size of the target object is within the allowable tolerance range;
[0146] Identify surface irregularities or machining errors;
[0147] Determine whether the spatial position of the target object meets the assembly requirements.
[0148] For target objects with abnormalities, the system marks them as defective and generates a detailed inspection report, including information such as deviation distribution and abnormal location.
[0149] Data visualization and output:
[0150] To facilitate subsequent quality control and analysis, the measurement and inspection results are output in three ways:
[0151] Point cloud image: contains the measurement points and key feature points of the target object, which is used to intuitively display the geometric shape of the target object.
[0152] Statistical data table: includes key information such as the measured size, position coordinates, deviation value, etc. of the target object.
[0153] Inspection Report: The system-generated inspection reports are output in a standard format for easy integration into industrial quality control processes.
[0154] This embodiment achieves comprehensive measurement and inspection of the three-dimensional features of the target object through the combination of deep learning models and point cloud geometry analysis, providing effective technical support for high-precision detection in complex industrial scenarios.
[0155] Step S5: Result feedback
[0156] In this embodiment, the measurement and inspection results of the target object are fed back to the industrial automation system in real time to achieve quality inspection and management of the automated production line. This step aims to improve the intelligence level and operation efficiency of the production line through real-time feedback and processing of data.
[0157] The measurement and inspection results are transmitted to the automation control system via an industrial communication bus (such as EtherCAT, PROFINET or MODBUSTCP / IP). The transmitted data includes:
[0158] Measurement data: including the position, size, shape and related deviation values of the target object;
[0159] Test mark: indicates the test result (pass / fail) and abnormal information of specific parts;
[0160] Statistical data: including real-time detection quantity, failure rate and trend analysis information.
[0161] In order to ensure the reliability and real-time performance of data transmission, the system uses a packet verification mechanism (such as CRC verification) to verify data integrity, and sets a low-latency transmission protocol to ensure that the measurement results can reach the control system within milliseconds.
[0162] After receiving the measurement results, the industrial automation system uses the preset judgment logic module to analyze and process the data. The judgment logic consists of the following rules:
[0163] Size determination: The system compares the measured size of the target object with the preset size range:
[0164]
[0165] Where L min and L max are the minimum and maximum lengths of the target object under the specification definition, and L is the length of the target object.
[0166] Shape matching judgment: The system uses the shape template matching algorithm to compare the measured shape of the target object with the standard shape template to determine whether the shape deviation is within the allowable range. The shape matching judgment formula is:
[0167]
[0168] Parameter Description:
[0169] ΔS: Shape deviation value, which is used to measure the average difference between the actual shape of the target object and the reference shape. The smaller the ΔS value, the closer the actual shape of the target object is to the reference shape.
[0170] N': The number of measurement points, used to normalize the shape matching. This ensures that the deviation value ΔS is the average of the deviations of all measurement points and is not affected by the number of measurement points.
[0171] S meas (i): The shape value of the i-th measurement point, which represents the actual shape parameter of the target object at the i-th position.
[0172] S ref (i): The shape value of the i-th reference point, which represents the ideal shape parameters of the standard shape template at the i-th position.
[0173] This formula calculates the measured shape S of the target object meas (i) With standard shape template S ref (i) The difference at each point is calculated, and the average of the differences at all points is taken to generate a shape deviation value ΔS. A smaller ΔS indicates that the shape of the target object is close to the template and meets the shape specification requirements; otherwise, it indicates that there is a large deviation in the shape of the target object.
[0174] Abnormal detection mark: For target objects that exceed the deviation range, the system generates abnormal marks and records the reasons for non-conformity (such as excessive size, inconsistent shape, etc.).
[0175] The system removes the target objects that are judged as "unqualified". The specific process is as follows:
[0176] Marking of rejected objects: The system triggers the marking device (such as an inkjet marking machine) installed on the conveyor belt to mark the surface of the target object for subsequent tracking by controlling the output signal.
[0177] Removal of unqualified objects: When the target object reaches the rejection station, the system removes the unqualified object from the production line through a pneumatic device, a robotic arm or a diverter device.
[0178] Data recording and feedback: The system records the measurement data and rejection information of non-conforming objects in the quality control database for subsequent tracking and statistical analysis.
[0179] After completing the real-time feedback, the system will also generate a comprehensive quality report for operators to monitor and optimize. The quality report includes:
[0180] Statistics of qualified rate and unqualified rate: display the quality status of the current production line;
[0181] Abnormal distribution and trend analysis: Use data visualization tools (such as bar charts and line charts) to display the reasons for non-conformity and change trends;
[0182] Equipment operating status: Monitor the operating status of measurement equipment and automation systems for timely adjustments.
[0183] Through the above steps, the measurement and inspection results in this embodiment can be fed back to the industrial automation system in real time and accurately, thereby realizing comprehensive monitoring and intelligent management of production line quality, and providing a guarantee for efficient operation in the field of industrial automation.
[0184] The above description is only a preferred embodiment of the present invention, so all equivalent changes or modifications made according to the structure, characteristics and principles described in the scope of the patent application of the present invention are included in the scope of the patent application of the present invention.
Claims
1. A method for improving 3D machine vision inspection and measurement accuracy based on a large model, characterized by: The following steps are involved: Step S1: collecting three-dimensional data of a target object by means of a multimodal sensor, wherein the multimodal sensor is used to obtain three-dimensional spatial information of the target object; Step S2: using an adaptive deep learning model to process the three-dimensional data to extract three-dimensional features of the target object, wherein the adaptive deep learning model can adjust model structure parameters in real time according to the three-dimensional data and environmental conditions; Step S3: Based on the three-dimensional features extracted by the adaptive deep learning model, the knowledge acquired by the model in other environments is transferred to the current target object detection task through transfer learning technology, and the transferred model is fine-tuned; Step S4: performing three-dimensional measurement and inspection on the three-dimensional features of the target object according to the fine-tuned deep learning model to generate position, size and shape data of the target object; Step S5: Feedback the three-dimensional measurement and inspection results to the industrial automation system.
2. A method for improving 3D machine vision inspection and measurement accuracy based on a large model as claimed in claim 1, characterized in that: The multimodal sensor includes at least one of a laser scanner, a stereo camera or a depth sensor.
3. A method for improving 3D machine vision inspection and measurement accuracy based on a large model as claimed in claim 2, characterized in that: The laser scanner is used to collect high-precision depth information of the target object, and the stereo camera is used to collect texture and surface details of the target object.
4. A method for improving 3D machine vision inspection and measurement accuracy based on a large model as claimed in claim 1, characterized in that: The adaptive deep learning model includes a dynamic adjustment module, which monitors environmental conditions through an environmental detector and adjusts the convolution kernel size, the number of network layers and the activation function type in real time according to the environmental conditions.
5. A method for improving 3D machine vision inspection and measurement accuracy based on a large model as claimed in claim 4, characterized in that: The environmental conditions monitored by the environmental detector include light intensity, background complexity, and reflectivity of the target object surface.
6. A method for improving 3D machine vision inspection and measurement accuracy based on a large model as claimed in claim 1, characterized in that: The three-dimensional measurement includes calculating the length, width and height of the target object through an outer envelope box, and the outer envelope box is generated by extreme points of the point cloud.
7. The method for improving 3D machine vision inspection and measurement accuracy based on a large model as claimed in claim 1, characterized in that: The three-dimensional inspection includes analyzing the shape characteristics of the target object and extracting the shape parameters of the target geometry through a least squares fitting algorithm.
8. The method for improving 3D machine vision inspection and measurement accuracy based on a large model as claimed in claim 1, characterized in that: The industrial automation system includes a quality detection module, which automatically determines whether the target object meets the preset specifications according to the measurement results, and performs a rejection operation on the target object that does not meet the requirements.
9. A method for improving 3D machine vision inspection and measurement accuracy based on a large model as claimed in claim 8, characterized in that: The rejection operation is performed by a pneumatic device, a robotic arm or a diversion device, and the detection results and rejection information of unqualified objects are recorded for subsequent tracking and statistical analysis.
Citation Information
Patent Citations
A mine card detection method and system based on 3D point cloud deep learning
CN109919145A
Low-illumination target detection method based on MSFAF-Net
CN117456330A
Multilevel context 3D target detection method based on surface bias
CN118470704A
Unmanned aerial vehicle traffic jam and personnel gathering detection method based on AI identification
CN119152413A
Intelligent labeling method and system for 3D point cloud data
CN119180868A
Cited By
Industrial application-oriented three-dimensional visual detection method, system and equipment
CN120510607A