A method for improving 3D machine vision inspection and measurement accuracy based on large models
Patent Information
- Application Number
- CN202510015736.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2045-01-06
AI Technical Summary
例如,激光扫描器在强光干扰下可能产生噪声,导致扫描数据的可靠性下降;立体相机在光线不足时难以提取准确的深度信息
[0025]This invention achieves efficient acquisition, accurate feature extraction, and intelligent measurement of 3D data of target objects by combining multimodal sensors, adaptive deep learning models, and transfer learning techniques. The adaptive model dynamically adjusts parameters to improve adaptability to complex environments, while transfer learning reduces the need for labeled data and enhances the model's generalization ability. Combined with real-time feedback from industrial automation systems, this invention significantly improves the inspection accuracy and measurement efficiency of 3D machine vision, meeting the high-precision inspection requirements in complex industrial scenarios.
Smart Images

Figure CN119941673B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of machine vision technology, and specifically to a method for improving the accuracy of 3D machine vision inspection and measurement based on large models. Background Technology
[0002] With the rapid development of industrial automation and intelligent manufacturing, 3D machine vision technology has become an indispensable tool in modern industrial production. In fields such as industrial quality inspection, assembly guidance, and production line monitoring, 3D machine vision technology enables precise measurement and inspection of targets by acquiring three-dimensional information about them. However, existing 3D machine vision technologies still face many limitations and technical bottlenecks in practical applications, hindering their widespread adoption and efficient application in complex industrial scenarios.
[0003] Currently, many 3D machine vision systems rely on a single type of sensor (such as a laser scanner or stereo camera) to acquire 3D data. While this single-sensor approach can provide stable data support in specific environments, it is often susceptible to interference from factors such as ambient lighting, target surface material, or texture deficiencies in variable industrial environments, thus affecting the integrity and reliability of data acquisition. For example, for highly reflective or textureless surfaces, traditional stereo vision methods may struggle to generate accurate 3D data, further limiting measurement precision.
[0004] At the data processing level, current 3D data processing methods largely rely on rule-based algorithms (such as feature matching and point cloud stitching), which are typically based on manually designed feature extraction rules. However, when the target object has complex geometry, smooth surface, blurred texture, or is in harsh environmental conditions such as low lighting and high dynamic range, these algorithms often lack adaptability and robustness, making it difficult to meet the high precision and stability requirements of industrial scenarios. Meanwhile, with the development of deep learning technology, some 3D machine vision systems have begun to adopt deep learning methods for feature extraction and object detection, but these data-driven methods usually require a large amount of labeled data. In industrial production environments, detection tasks are diverse, target objects are varied, and production environments are complex, making the acquisition of labeled data extremely costly. Furthermore, the significant differences in equipment configuration and environmental conditions across different production lines make it difficult to directly transfer and apply existing trained models, increasing the difficulty and cost of model deployment.
[0005] Industrial production lines place higher demands on the real-time performance and accuracy of 3D machine vision systems. However, traditional algorithms and some deep learning models have high computational complexity, making it difficult for the system's processing speed to meet the high-speed operation requirements of production lines when dealing with large-scale point cloud data or high-resolution images. This performance bottleneck is particularly pronounced in real-time application environments, further limiting its widespread application in industrial scenarios.
[0006] Furthermore, the development of multimodal data fusion technology is still immature. Industrial applications often face multiple complex scenarios, and relying solely on a single type of 3D data cannot fully reflect the characteristics of the target object. Although multimodal sensor fusion technology has made some progress in recent years, in industrial practice, different sensors have different data acquisition cycles and significant differences in accuracy. How to efficiently fuse these heterogeneous data remains a technical challenge. At the same time, the redundancy and consistency issues of multimodal data also place higher demands on the realization of high-precision detection.
[0007] Traditional 3D machine vision systems typically employ static design, meaning they operate within a specific scene or target area with fixed parameters. When the type of target object or the production environment changes, existing systems often require reconfiguration or retraining, lacking the ability to dynamically adjust. This static characteristic not only limits the system's versatility but also poses a significant challenge to its ability to adapt to new environments.
[0008] In complex industrial environments, external interferences such as changes in lighting, dust, and vibration often affect the performance of 3D machine vision systems. For example, laser scanners may generate noise under strong light interference, leading to a decrease in the reliability of scan data; stereo cameras struggle to extract accurate depth information in low light conditions. These problems severely limit the stability of 3D machine vision systems, making them face more uncertainties in practical industrial applications.
[0009] In summary, existing 3D machine vision technologies face technical bottlenecks in areas such as data acquisition integrity, algorithm adaptability, model transferability, balance between real-time performance and accuracy, multimodal data fusion, and system dynamic adjustment capabilities. New technical solutions are urgently needed to overcome these limitations and better meet the complex needs of industrial scenarios. Summary of the Invention
[0010] To address the aforementioned problems in the prior art, this invention proposes a method for improving the accuracy of 3D machine vision inspection and measurement based on large models, characterized by the following steps:
[0011] Step S1: Collect three-dimensional data of the target object using a multimodal sensor, wherein the multimodal sensor is used to acquire the three-dimensional spatial information of the target object;
[0012] Step S2: Use an adaptive deep learning model to process the 3D data to extract the 3D features of the target object, wherein the adaptive deep learning model can adjust the model structure parameters in real time according to the 3D data and environmental conditions;
[0013] Step S3: Based on the 3D features extracted by the adaptive deep learning model, the knowledge obtained by the model in other environments is transferred to the current target object detection task through transfer learning technology, and the transferred model is fine-tuned.
[0014] Step S4: Based on the fine-tuned deep learning model, perform three-dimensional measurement and inspection on the three-dimensional features of the target object to generate the position, size and shape data of the target object;
[0015] Step S5: Feed back the three-dimensional measurement and inspection results to the industrial automation system.
[0016] The multimodal sensor includes at least one laser scanner, stereo camera, or depth sensor.
[0017] The laser scanner is used to acquire high-precision depth information of the target object, and the stereo camera is used to acquire the texture and surface details of the target object.
[0018] The adaptive deep learning model includes a dynamic adjustment module, which monitors environmental conditions through an environmental detector and adjusts the convolutional kernel size, the number of network layers, and the activation function type in real time according to the environmental conditions.
[0019] The environmental conditions monitored by the environmental detector include light intensity, background complexity, and the reflectivity of the target object's surface.
[0020] The three-dimensional measurement includes calculating the length, width, and height of the target object using an outer envelope box generated from point cloud extrema.
[0021] The three-dimensional inspection includes analyzing the shape characteristics of the target object and extracting the shape parameters of the target geometry using a least-squares fitting algorithm.
[0022] The industrial automation system includes a quality inspection module, which automatically determines whether the target object meets the preset specifications based on the measurement results, and performs a rejection operation on the target object that does not meet the requirements.
[0023] The rejection operation is performed by a pneumatic device, a robotic arm, or a diversion device, and the detection results and rejection information of the non-conforming objects are recorded for subsequent tracking and statistical analysis.
[0024] Beneficial effects:
[0025] This invention achieves efficient acquisition, accurate feature extraction, and intelligent measurement of 3D data of target objects by combining multimodal sensors, adaptive deep learning models, and transfer learning techniques. The adaptive model dynamically adjusts parameters to improve adaptability to complex environments, while transfer learning reduces the need for labeled data and enhances the model's generalization ability. Combined with real-time feedback from industrial automation systems, this invention significantly improves the inspection accuracy and measurement efficiency of 3D machine vision, meeting the high-precision inspection requirements in complex industrial scenarios. Attached Figure Description
[0026] The accompanying drawings, which are provided to further illustrate the invention and form part of this application, are not intended to unduly limit the invention. In the drawings:
[0027] Figure 1 A flowchart of the method of the present invention is shown. Detailed Implementation
[0028] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. The illustrative embodiments and descriptions are only used to explain the present invention and are not intended to limit the present invention.
[0029] This invention provides a method for improving the accuracy of 3D machine vision inspection and measurement based on large models. By employing multimodal sensors, adaptive deep learning models, and transfer learning techniques, it achieves high-precision 3D measurement and inspection of target objects, and applies the measurement results to industrial automation systems. The method is described in detail below with reference to specific embodiments.
[0030] Step S1: 3D Data Acquisition
[0031] In this embodiment, a multimodal sensor system is used to comprehensively acquire data about the target object in order to obtain high-precision three-dimensional data. The multimodal sensor system includes a laser scanner and a stereo camera, which work synchronously through a preset time synchronization module to ensure the accuracy and consistency of data acquisition.
[0032] The laser scanner is mounted on a stable stand and performs multi-angle scanning of the target object. Specifically, the laser scanner uses its built-in laser beam projection system to project a laser beam onto the surface of the target object and receives the reflected light signal through a built-in optical sensor to generate depth data of the target object. The depth data is converted into a depth map by a processing unit, with an accuracy range of 0.01 mm to 0.1 mm. The frequency and sampling rate of the laser beam can be adjusted according to specific application requirements to meet different accuracy requirements. The operating parameters of the laser scanner include:
[0033] Laser beam wavelength: 650 nanometers;
[0034] Scanning angle range: ±60 degrees;
[0035] Sampling frequency: 1000 times / second.
[0036] The control signals for the laser scanner are generated by the data acquisition module and sent to the laser scanner drive unit via an industrial protocol to ensure precise control of the scanning path and frequency.
[0037] A stereo camera is mounted on a bracket parallel to the laser scanner to acquire texture and surface feature information of the target object. The stereo camera includes two high-resolution imaging units that capture images of the target object from different angles, and generates 3D texture data associated with the target object using a stereo matching algorithm. The operating parameters of the stereo camera include:
[0038] Image resolution: 1920×1080 pixels;
[0039] Frame rate: 60 frames per second;
[0040] Lens focal length: 50 mm.
[0041] The camera's exposure time and gain parameters are automatically adjusted according to the shooting environment. For example, in a high-reflection environment, the camera reduces the exposure time through its built-in exposure adjustment module to avoid overexposure; in a low-light environment, the gain parameter is automatically increased to enhance texture details.
[0042] To ensure the timing consistency of data acquired by the laser scanner and stereo camera, the system is equipped with a hardware-based time synchronization module. This module uses a high-precision clock chip (nanosecond-level accuracy) as the master clock source, driving the synchronous startup of the laser scanner and stereo camera through a synchronization pulse signal (trigger signal). The synchronization pulse signal is generated by the FPGA module, with a period of 1 millisecond, and precise allocation of the trigger signal is achieved through clock frequency division. The synchronization module includes the following key components:
[0043] Master clock chip: Built-in high-precision oscillator with frequency stability of ±5ppm;
[0044] FPGA control unit: generates synchronous trigger signals to control the data acquisition cycle;
[0045] Synchronization interface: Uses standard LVDS signal to connect to the laser scanner and stereo camera.
[0046] After the laser scanner and stereo camera complete the data acquisition, the generated depth map and texture map are transmitted to the data processing unit via a high-speed data bus. In the data processing unit, a spatial alignment algorithm is used to fuse the depth map and texture map, ensuring that their correspondence in 3D space is consistent. The fused 3D dataset is output as a point cloud in standard PLY (Polygon File Format) or OBJ format, containing the coordinates (X, Y, Z) and color (R, G, B) information for each point.
[0047] The laser scanner provides high-precision depth data, while the stereo camera supplements the target surface texture information, ensuring the comprehensiveness of the acquired data. Both the laser scanner and the stereo camera are environmentally adaptive, dynamically adjusting acquisition parameters according to the actual acquisition scenario. A hardware-based time synchronization module provides nanosecond-level synchronization accuracy, ensuring the consistency of multimodal data across the time dimension. The fused 3D data adopts a universal point cloud format for easy subsequent processing and analysis.
[0048] Through the above technical means, this embodiment achieves comprehensive acquisition of high-precision three-dimensional data of the target object, laying the foundation for subsequent feature extraction, inspection and measurement.
[0049] Step S2: 3D Feature Extraction
[0050] In this embodiment, the acquired 3D data is input into an adaptive deep learning model for processing to extract the 3D features of the target object. The adaptive deep learning model is built on a Convolutional Neural Network (CNN) architecture and can automatically learn and extract effective features from the 3D data, thereby providing basic data support for subsequent inspection and measurement steps.
[0051] The adaptive deep learning model employs a multi-layer convolutional network structure, including an input layer, several convolutional layers, pooling layers, and fully connected layers. The input layer receives preprocessed 3D data (such as depth maps and texture maps) and transforms it into a format suitable for model computation. The convolutional layers extract spatial features of the target object, the pooling layers are used for dimensionality reduction to decrease computational complexity, and the fully connected layers map the extracted features into fixed-length feature vectors for use in subsequent steps.
[0052] Specifically, the key layer configuration of the model is as follows:
[0053] The convolutional layers use 3×3 or 5×5 kernels with a stride of 1 and ReLU (Rectified Linear Unit) activation function.
[0054] The pooling layer uses a 2×2 max pooling operation to reduce the feature map size while preserving the main features.
[0055] The dimension of the feature vector output by the fully connected layer is set according to the specific application requirements, and is commonly set to 256 dimensions or 512 dimensions.
[0056] To enhance the model's environmental adaptability, the adaptive deep learning model is equipped with a dynamic tuning module. This module consists of an environment detector and a parameter tuning unit, with the following specific functions:
[0057] The environmental detector monitors environmental conditions in real time during data acquisition using sensors, including parameters such as light intensity, surface reflectivity of the target object, and background complexity. For example, by acquiring the light distribution histogram of the depth map and the contrast value of the texture map, it calculates indices of ambient lighting uniformity and target surface characteristics. The detected environmental information is transmitted to the parameter adjustment unit in numerical form.
[0058] The parameter adjustment unit dynamically adjusts the key parameters of the adaptive deep learning model according to preset rules, mainly including convolutional kernel size, number of network layers, and activation function type. The adjustment logic is based on the following rules:
[0059] Low light conditions: Increase the kernel size (e.g., from 3×3 to 5×5) to capture more details, and increase the number of network layers to improve feature extraction capabilities.
[0060] High reflectivity surface: Adjust the activation function to Sigmoid to enhance the response to subtle features, while reducing pooling operations to retain more local information.
[0061] For textured environments: reduce the kernel size (e.g., from 5×5 to 3×3) and reduce the number of network layers to improve processing speed.
[0062] The parameter adjustment is based on a dynamic parameter control algorithm, and the specific formula is as follows:
[0063] K new =K default +ΔK·f(E)
[0064] in:
[0065] K new This refers to the adjusted kernel size.
[0066] K default This is the default kernel size;
[0067] ΔK is the adjustment increment, which typically ranges from 1 to 2;
[0068] f(E) is an environmental characteristic function, calculated based on environmental parameters such as light intensity and surface characteristics, and is defined as follows:
[0069]
[0070] Where E light This represents the light intensity index, with a value ranging from 0 to 1.
[0071] Feature extraction process flow:
[0072] Input 3D data, including depth maps and texture maps, is normalized and then passed to the input layer;
[0073] The environmental detector analyzes the environmental characteristics of the input data in real time and transmits the detection results to the parameter adjustment unit;
[0074] The parameter adjustment unit dynamically adjusts the convolutional kernel size, the number of network layers, and the activation function type to generate an optimized model structure;
[0075] The adaptive deep learning model extracts three-dimensional features based on the optimized parameter structure and generates feature vectors.
[0076] The extracted feature vectors are output to subsequent steps.
[0077] Step S3: Transfer Learning and Fine-tuning
[0078] In this embodiment, in order to reduce the need for labeled data and improve the adaptability and performance of the model in the new environment, transfer learning technology is used to transfer the knowledge of the pre-trained deep learning model to the target production environment, and the model is fine-tuned with a small amount of labeled data to adapt it to the specific needs of the current production line and the characteristics of the target object.
[0079] This embodiment selects a pre-trained deep learning model as the base model from an industry-standard dataset. This model is built on a convolutional neural network (CNN) architecture and has been thoroughly trained on general 3D object detection tasks, enabling it to extract high-quality 3D features. The specific configuration of the base model is as follows:
[0080] The input layer supports multimodal data input, including depth maps and texture maps;
[0081] The feature extraction network contains multiple convolutional layers and pooling layers;
[0082] The output layer contains target classification nodes and location regression nodes.
[0083] The training data for the basic model includes a wide range of industrial objects (such as parts, device surfaces, etc.) and their 3D data, covering common geometric shapes and surface material properties, providing a strong general knowledge foundation for transfer learning.
[0084] Transfer learning process:
[0085] At the start of transfer learning, the base feature extraction layers (low-level convolutional networks) of the pre-trained model are frozen, retaining their original weights to maintain general 3D feature extraction capabilities. These features are highly universal for most target objects and do not require readjustment during the transfer process.
[0086] Add a new task-specific layer on top of the base model to adapt to the detection requirements of the target production environment. For example, add shape parameter prediction nodes to the target location regression, or adjust the number and types of classification nodes according to the target characteristics.
[0087] For the new task layer, initial weights are generated using random initialization, and initial training is performed using a small amount of labeled data from the target production environment. This process optimizes the parameters of the new task layer through backpropagation, gradually adapting it to the new environment.
[0088] After completing the initial training of the new task layers, the parameters of some basic feature extraction layers (such as the later convolutional network layers) are unfrozen, and the entire model is jointly fine-tuned. During the fine-tuning process, the learning rate is dynamically adjusted according to the characteristics of the target environment. Typically, a smaller learning rate (such as 0.0001) is chosen for the basic feature extraction layers, while a larger learning rate (such as 0.001) is used for the new task layers to ensure that the model adapts efficiently to the new task.
[0089] The fine-tuning process employs the Mini-Batch Gradient Descent (MBGD) method to optimize the objective function and minimize the target detection error. The specific objective function L′ is as follows:
[0090] L′=α·L cls +β·L reg
[0091] in:
[0092] L cls For classification loss, the cross-entropy loss function is used;
[0093] L reg For location regression loss, use the L2 loss function;
[0094] α and β are weighting coefficients used to balance the influence of classification and regression tasks, and are usually set to 1:1.
[0095] During fine-tuning, the system dynamically samples labeled data from the new environment and combines this with transfer learning techniques to improve the model's ability to specifically detect the target environment. For example, when detecting defects on metal surfaces, the model can more accurately identify dents and scratches through fine-tuning, while maintaining universality for complex geometries.
[0096] After fine-tuning, the performance of the fine-tuned model in the target environment was evaluated through cross-validation, including object detection precision, recall, and F1 score. Workpieces in the target production environment were used as test objects to verify its application effect on a real production line. Tests show that transfer learning combined with the fine-tuned model can significantly improve the detection accuracy of target objects while reducing the dependence on large-scale labeled data.
[0097] This embodiment successfully achieves rapid adaptation of the basic model to the target environment through transfer learning and fine-tuning techniques, providing an effective solution for high-precision target detection and measurement in complex industrial production environments.
[0098] Step S4: Three-dimensional measurement and inspection
[0099] In this embodiment, a deep learning model that has undergone transfer learning and fine-tuning optimization is used to perform three-dimensional measurement and inspection of the target object, aiming to generate position, size and shape data of the target object, and ensure comprehensive coverage and accurate extraction of key features of complex-shaped objects.
[0100] 3D data input and preprocessing:
[0101] The acquired depth and texture maps are fed as input data into the optimized deep learning model. To improve the accuracy of measurement and inspection, the input data is first preprocessed, including the following steps:
[0102] Depth map normalization: Maps depth values in the depth map to the range of [0,1] to eliminate measurement scale differences between different sensors.
[0103] Texture map enhancement: Enhance the contrast of texture maps through histogram equalization to highlight the surface features of the target object and reduce data loss under high reflectivity or low lighting conditions.
[0104] Coordinate Alignment and Fusion: A spatial alignment algorithm is used to unify the coordinate systems of the depth map and texture map, ensuring data consistency and accuracy. The aligned data is then used to generate a 3D point cloud of the target object through a point cloud fusion algorithm.
[0105] 3D Feature Extraction and Key Feature Recognition:
[0106] The optimized deep learning model processes point cloud data to extract the 3D features of the target object and identify key geometric features. The model's feature extraction network includes multiple convolutional operations, enabling it to extract the object's spatial shape characteristics and local geometric features, such as curvature, edges, and vertices. The key feature identification process is based on the following steps:
[0107] Curvature Calculation: Key surface regions of an object are identified using the principal curvature calculation formula for the local neighborhood of the point cloud. The specific formula is as follows:
[0108]
[0109] This formula is used to calculate the principal curvature of a local region in a point cloud, which is used to identify key curved surface regions of a target object. Principal curvature is one of the important features characterizing the local geometry of a point cloud, reflecting the degree of curvature of an object's surface in a certain direction.
[0110] The formula parameters are explained as follows:
[0111] C: Represents the curvature value, which is the result of the formula. The value of C ranges between [0,1] and is used to measure the surface characteristics of the point cloud neighborhood.
[0112] When C is close to 0, it indicates that the surface of the region is relatively flat;
[0113] When C is close to 1, it indicates that the region has sharp features, such as edges or vertices.
[0114] λ max : Represents the maximum eigenvalue (maximum principal curvature) in the principal direction of a local neighborhood of a point cloud.
[0115] It is the largest eigenvalue in the covariance matrix of the neighborhood point cloud, and usually represents the intensity of change of the local surface in the direction of maximum curvature.
[0116] λ min : Represents the minimum eigenvalue (minimum principal curvature) of the principal direction in the local neighborhood of the point cloud.
[0117] It is the smallest eigenvalue in the covariance matrix of the neighborhood point cloud, and usually reflects the intensity of the change in the local surface in the direction of minimum curvature.
[0118] Neighborhood point cloud: refers to a local point cloud set centered on the target point, defined by a radius r or the number of points k. The principal direction eigenvalue λ is determined by calculating the covariance matrix of these neighboring points. max and λ min .
[0119] This formula uses the principal curvature ratio C to quantitatively represent the degree of curvature of an object's surface.
[0120] In point cloud processing, the curvature value C is used to distinguish between flat areas, curved areas, and sharp areas (such as edges and vertices) of an object.
[0121] A larger curvature value C indicates that the point is located at a significant feature position on the object's surface, making it suitable as a key feature point for further processing.
[0122] This formula, combined with point cloud local feature calculation, can effectively support the 3D feature extraction and key point recognition of complex-shaped objects.
[0123] Edge detection: Gradient-based edge detection algorithms are used to identify the boundary regions of target objects, thereby extracting their contour features.
[0124] Key point extraction: Calculate key points, such as cusps or intersections, based on the geometric discontinuities in the point cloud for subsequent measurement and analysis.
[0125] 3D Measurement and Data Output:
[0126] When measuring the position, size, and shape of a target object, the system generates high-precision measurement data based on the spatial distribution and geometric characteristics of the point cloud. The specific process is as follows:
[0127] Position measurement: Calculate the 3D position of the target object using the centroid of the point cloud. Centroid coordinates (X... c ,Y c Z c The formula for calculating ) is:
[0128]
[0129] Parameter description:
[0130] X c ,Y c Z c : These represent the centroid coordinates of the target object in three-dimensional space. The centroid is the geometric center of the object's point cloud and represents the position of the target object.
[0131] N: The total number of points in the point cloud. It is used to calculate the average value of all points in the point cloud, ensuring that the centroid position is closely related to the distribution of all points.
[0132] x i ,y i ,z i : The coordinates of the i-th point in the three-dimensional coordinate system. The position of each point in the point cloud is represented by these coordinates.
[0133] This formula generates the centroid coordinates (X, Y) of the target object by calculating the average of the three-dimensional coordinates of all points. c ,Y c Z c The center of mass reflects the position of the target object in three-dimensional space and is the key to the overall positioning of the target object.
[0134] Dimensional Measurement: The length, width, and height of the target object are calculated using the bounding box of the point cloud. The bounding box is generated from the extreme points of the point cloud, using the following formula:
[0135] L = X max -X min W = Y max -Y min H = Z max -Z min
[0136] Parameter description:
[0137] L, W, H: represent the length, width, and height of the target object, respectively, and are a three-dimensional representation of the target object's dimensions.
[0138] X max ,X min The maximum and minimum values of the point cloud in the X direction correspond to the two farthest points of the point cloud along the X-axis. L is calculated using the difference between these two extreme points, representing the size of the target object in the X direction.
[0139] Y max ,Y min : The maximum and minimum values of the point cloud in the Y direction, corresponding to the two farthest points of the point cloud distribution along the Y-axis. W is calculated using the difference between these two extreme points, representing the size of the target object in the Y direction.
[0140] Z max Z min The maximum and minimum values of the point cloud in the Z direction correspond to the two farthest points of the point cloud along the Z-axis. H is calculated using the difference between these two extreme points, representing the size of the target object in the Z direction.
[0141] This formula generates the 3D dimensions of a target object by calculating the difference between extreme points of the point cloud on the 3D coordinate axes. It is used to quickly approximate the volume and spatial extent of a target object.
[0142] Shape analysis: Extract the shape characteristics of the target object, such as a plane, cylinder, or sphere, based on a point cloud fitting algorithm. Generate a mathematical model of the target geometry using a least-squares fitting method and calculate its specific parameters (such as radius, tilt angle, etc.).
[0143] Comprehensive analysis and inspection:
[0144] After completing the measurement, the system compares and analyzes the actual measurement data of the target object with preset standards to identify potential deviations or defects. For example:
[0145] Check whether the dimensions of the target object are within the allowable tolerances;
[0146] Identify surface irregularities or processing errors;
[0147] Determine whether the spatial position of the target object meets the assembly requirements.
[0148] For target objects that exhibit abnormalities, the system marks them as non-conforming products and generates a detailed inspection report, including information such as deviation distribution and abnormal location.
[0149] Data visualization and output:
[0150] To facilitate subsequent quality control and analysis, measurement and inspection results are output in three ways:
[0151] Point cloud map: Contains measurement points and key feature points of the target object, used to visually display the geometry of the target object.
[0152] Statistical data table: Includes key information such as the measured dimensions, position coordinates, and deviation values of the target object.
[0153] Inspection Report: The inspection reports generated by the system are output in a standard format, which facilitates integration into industrial quality control processes.
[0154] This embodiment combines deep learning models with point cloud geometric analysis to achieve comprehensive measurement and inspection of the three-dimensional features of target objects, providing effective technical support for high-precision detection in complex industrial scenarios.
[0155] Step S5: Result Feedback
[0156] In this embodiment, the measurement and inspection results of the target object are fed back to the industrial automation system in real time to achieve quality inspection and management of the automated production line. This step aims to improve the intelligence level and operational efficiency of the production line through real-time data feedback and processing.
[0157] Measurement and inspection results are transmitted to the automation control system via industrial communication buses (such as EtherCAT, PROFINET, or MODBUS TCP / IP). The transmitted data includes:
[0158] Measurement data: including the position, size, shape, and related deviation values of the target object;
[0159] Inspection markings: Indicate the inspection result (pass / fail) and abnormal information of specific parts;
[0160] Statistical data: including real-time detection volume, failure rate, and trend analysis information.
[0161] To ensure the reliability and real-time performance of data transmission, the system employs a data packet verification mechanism (such as CRC check) to verify data integrity, and sets up a low-latency transmission protocol to ensure that measurement results can reach the control system within milliseconds.
[0162] After receiving the measurement results, the industrial automation system analyzes and processes the data using a pre-set judgment logic module. The judgment logic consists of the following rules:
[0163] Size and Specification Judgment: The system compares the measured dimensions of the target object with a preset specification range.
[0164]
[0165] Where L min and L max These are the minimum and maximum lengths of the target object as defined in the specifications, where L is the length of the target object.
[0166] Shape matching judgment: The system compares the measured shape of the target object with a standard shape template using a shape template matching algorithm to determine whether the shape deviation is within the allowable range. The shape matching judgment formula is as follows:
[0167]
[0168] Parameter description:
[0169] ΔS: Shape deviation value, used to measure the average difference between the actual shape of a target object and a reference shape. The smaller the value of ΔS, the closer the actual shape of the target object is to the reference shape.
[0170] N': The number of measurement points used for normalization of shape matching. This ensures that the deviation value ΔS is the average of the deviations across all measurement points, unaffected by the number of measurement points.
[0171] S meas (i): The shape value of the i-th measurement point, representing the actual shape parameters of the target object at the i-th position.
[0172] S ref (i): The shape value of the i-th reference point, representing the ideal shape parameters of the standard shape template at the i-th position.
[0173] This formula calculates the measured shape S of the target object. meas (i) with standard shape template S ref (i) Calculate the difference at each point and average the differences across all points to generate a shape deviation value ΔS. A smaller ΔS indicates that the shape of the target object is close to the template and meets the shape specifications; conversely, a larger ΔS indicates that the shape of the target object has a large deviation.
[0174] Anomaly detection marker: For target objects that exceed the deviation range, the system generates an anomaly marker and records the reason for non-compliance (such as excessive size, incorrect shape, etc.).
[0175] The system removes target objects deemed "unacceptable". The specific process is as follows:
[0176] Marking defective objects: The system triggers a marking device (such as an inkjet marker) installed on the conveyor belt to mark the surface of the target object by controlling the output signal, so as to facilitate subsequent tracking.
[0177] Removing non-conforming objects: When a target object arrives at the rejection station, the system removes the non-conforming object from the production line using a pneumatic device, robotic arm, or diversion device.
[0178] Data recording and feedback: The system records the measurement data and rejection information of non-conforming objects in the quality control database for subsequent tracking and statistical analysis.
[0179] After providing real-time feedback, the system will also generate a comprehensive quality report to facilitate monitoring and optimization by operators. The quality report includes:
[0180] Pass rate and fail rate statistics: Displays the current quality status of the production line;
[0181] Anomaly Distribution and Trend Analysis: Using data visualization tools (such as bar charts and line charts) to display the causes of non-compliance and trends of change;
[0182] Equipment operating status: Monitor the operating status of measuring equipment and automation systems to make timely adjustments.
[0183] Through the above steps, the measurement and inspection results in this embodiment can be fed back to the industrial automation system in real time and accurately, thereby realizing comprehensive monitoring and intelligent management of production line quality and providing a guarantee for efficient operation in the field of industrial automation.
[0184] The above description is only a preferred embodiment of the present invention. Therefore, all equivalent changes or modifications made to the structure, features and principles described in the claims of this patent application are included in the scope of this patent application.
Claims
1. A method for improving the accuracy of 3D machine vision inspection and measurement based on large models, characterized in that: Includes the following steps: Step S1: Collect three-dimensional data of the target object using a multimodal sensor, wherein the multimodal sensor is used to acquire the three-dimensional spatial information of the target object; Step S2: Use an adaptive deep learning model to process the 3D data to extract the 3D features of the target object, wherein the adaptive deep learning model can adjust the model structure parameters in real time according to the 3D data and environmental conditions; Step S3: Based on the 3D features extracted by the adaptive deep learning model, the knowledge gained by the model in other environments is transferred to the current target object detection task using transfer learning technology, and the transferred model is fine-tuned. The adaptive deep learning model includes a dynamic adjustment module, which monitors environmental conditions through an environmental detector and adjusts the convolutional kernel size, number of network layers, and activation function type in real time according to the environmental conditions. The adjustment logic is based on the following rules: Low light conditions: Increase the kernel size and the number of network layers; High reflectivity surface: Adjust the activation function to Sigmoid, and reduce pooling operations; Richly textured environments: Reduce convolutional kernel size and thus reduce the number of network layers; The parameter adjustment is based on a dynamic parameter control algorithm, and the specific formula is as follows: in: This refers to the adjusted kernel size. This is the default kernel size; To adjust the increment; The environmental characteristic function is defined as follows: in This represents the light intensity index, with a value ranging from 0 to 1; Step S4: Based on the fine-tuned deep learning model, perform three-dimensional measurement and inspection on the three-dimensional features of the target object to generate the position, size and shape data of the target object; the three-dimensional inspection includes analyzing the shape characteristics of the target object and extracting the geometric shape parameters of the target object through the least squares fitting algorithm; Step S5: Feed back the three-dimensional measurement and inspection results to the industrial automation system.
2. The method for improving the accuracy of 3D machine vision inspection and measurement based on a large model as described in claim 1, characterized in that: The multimodal sensor includes at least one laser scanner, stereo camera, or depth sensor.
3. The method for improving the accuracy of 3D machine vision inspection and measurement based on a large model as described in claim 2, characterized in that: The laser scanner is used to acquire high-precision depth information of the target object, and the stereo camera is used to acquire the texture and surface details of the target object.
4. The method for improving the accuracy of 3D machine vision inspection and measurement based on a large model as described in claim 1, characterized in that: The environmental conditions monitored by the environmental detector include light intensity, background complexity, and the reflectivity of the target object's surface.
5. The method for improving the accuracy of 3D machine vision inspection and measurement based on a large model as described in claim 1, characterized in that: The three-dimensional measurement includes calculating the length, width, and height of the target object using an outer envelope box generated from point cloud extrema.
6. The method for improving the accuracy of 3D machine vision inspection and measurement based on a large model as described in claim 1, characterized in that: The industrial automation system includes a quality inspection module, which automatically determines whether the target object meets the preset specifications based on the measurement results, and performs a rejection operation on the target object that does not meet the requirements.
7. The method for improving the accuracy of 3D machine vision inspection and measurement based on a large model as described in claim 6, characterized in that: The rejection operation is performed by a pneumatic device, a robotic arm, or a diversion device, and the detection results and rejection information of the non-conforming objects are recorded for subsequent tracking and statistical analysis.
Citation Information
Patent Citations
A mine card detection method and system based on 3D point cloud deep learning
CN109919145A
Multilevel context 3D target detection method based on surface bias
CN118470704A
Unmanned aerial vehicle traffic jam and personnel gathering detection method based on AI identification
CN119152413A