Robot vision system construction and control method thereof

By collecting data through multispectral cameras and inertial measurement units, combined with image processing and reinforcement learning, the robot vision system can achieve adaptive compensation and path correction under complex lighting conditions, solving the problems of target recognition deviation and grasping failure, and improving positioning accuracy and operation success rate.

CN120755881APending Publication Date: 2025-10-10SUZHOU YOULIAN WEISHI TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511089638.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Under the complex lighting conditions of industrial sites, it is difficult for robot vision systems to distinguish abnormal light reflections and motion blur in real time, resulting in deviations in target recognition results, affecting positioning accuracy and grasping success rate.

Method used

A multispectral camera array and inertial measurement unit are used to collect environmental data, and the images are processed by combining adaptive median filtering, contrast enhancement and geometric correction. A convolutional neural network is used to extract visual features, and a reinforcement learning decision model is combined to generate robot motion control instructions. The control strategy is dynamically adjusted to optimize the data to achieve adaptive compensation and path correction.

Benefits of technology

Under complex lighting conditions, the stability of visual data acquisition and the accuracy of target positioning are improved, the trajectory error of vision guidance is reduced, the risk of grasping failure and collision is avoided, and the safety and success rate of robot operations are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120755881A_ABST
    Figure CN120755881A_ABST
Patent Text Reader

Abstract

The invention discloses a robot vision system construction and control method, and the method comprises the following steps: collecting environment image data and robot motion state data, carrying out the real-time preprocessing of the environment image data, generating standardized image data, and carrying out the real-time monitoring of the environment through setting a dynamic environment sensing module. When real-time grabbing operation guided by robot vision is carried out, an environment light interference real-time judgment mechanism is established, differential compensation strategies are set for different working condition scenes, the stability of visual data collection under the complex illumination condition is guaranteed, and meanwhile compensated light parameters are synchronized to an image processing unit; the problem of feature recognition distortion caused by abnormal environmental reflection can be reduced in real time, the accuracy of robot target positioning is ensured, the track error of visual guidance is reduced, and the completeness of a high-precision assembly task is ensured when it is detected that the in-place posture offset exceeds a safety threshold value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a robot vision system construction and a control method thereof. Background Art

[0002] Artificial intelligence is a new technological science that studies and develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence.

[0003] The application of robotic vision systems in intelligent grasping and precise positioning tasks in the field of industrial automation is a comprehensive system that integrates machine vision, motion control and artificial intelligence technologies, aiming to improve the accuracy and response speed of robotic operations. This is one of the core functional modules of the system, responsible for real-time analysis of environmental information through visual sensors. The high-frame rate image acquisition unit can obtain continuous images of dynamic scenes to ensure accurate identification of the shape and posture of target objects under complex working conditions. To ensure the consistency and reliability of visual data, the system needs to be configured with an adaptive ambient light compensation mechanism to ensure the stable operation of the robotic vision system.

[0004] At present, due to the dynamic lighting changes and equipment movement interference in the industrial field environment, when performing real-time grasping operations guided by robot vision, the deployed visual sensors cannot determine in real time whether the current image frame has abnormal lighting reflection and motion blur when capturing the surface texture of the target object. When the reflective area causes the feature points to be distorted, it will cause deviations in the target recognition results, making it difficult to ensure the accuracy of the positioning results.

[0005] Therefore, a robot vision system construction and control method are proposed to solve the above problems. Summary of the Invention

[0006] In view of the deficiencies in the prior art, the present invention provides a robot vision system construction and a control method thereof, which solves the problems raised in the above background technology.

[0007] To achieve the above objectives, the present invention provides the following technical solutions: a robot vision system construction and control method, the method comprising the following steps: S1, collect environmental image data and robot motion state data; S2. Preprocess the environmental image data in real time to generate standardized image data; S3. Extracting visual feature parameters based on the standardized image data to generate a visual feature descriptor; S4, performing target object recognition and positioning processing according to the visual feature descriptor and the robot motion state data to generate target recognition result data; S5, generating robot motion control instructions based on the target recognition result data and a preset control strategy; S6, executing the robot motion control instruction to drive the robot actuator to complete the action; S7. Dynamically adjust control strategy parameters based on action execution feedback data to generate adaptive control optimization data.

[0008] Preferably, said S1 comprises the following steps: S11, collecting visible light and infrared image data in the environment through a multispectral camera array; S12, collecting acceleration, angular velocity and attitude data of the robot through an inertial measurement unit; S13, fusing the visible light and infrared image data with the posture data to generate a spatiotemporally synchronized environmental image data set; S14: performing timestamp alignment processing on the environmental image dataset to ensure consistency of data acquisition frequency.

[0009] Preferably, said S2 comprises the following steps: S21, performing denoising filtering on the environmental image data, using an adaptive median filter to eliminate impulse noise; S22, performing contrast enhancement processing on the denoised image data, using a limited contrast adaptive histogram equalization algorithm; S23, performing geometric distortion correction on the enhanced image data, and performing perspective transformation based on the camera intrinsic parameter model; S24 , normalizing the corrected image data to a uniform resolution to generate the standardized image data.

[0010] Preferably, said S3 comprises the following steps: S31, using a convolutional neural network model to extract spatial features of the standardized image data and generate a primary feature map; S32. Generate feature descriptors through multi-scale pooling and perform feature fusion: , in, is the normalized fusion feature vector, is the spatial eigenvector, is the frequency domain eigenvector, is the weighting coefficient, , the default threshold ; S33, fusing the multi-resolution feature descriptor with the robot motion state data to construct a spatiotemporal joint feature vector; S34. Reduce the dimension of the spatiotemporal joint feature vector by principal component analysis to generate the visual feature descriptor.

[0011] Preferably, said S4 comprises the following steps: S41, inputting the visual feature descriptor into a YOLOv7 target detection model to identify dynamic and static targets in the environment; S42, calculating the 3D position coordinates of the target object based on the binocular vision triangulation principle; S43, combining the robot posture data, converting the target position to the robot coordinate system; S44: Generate target recognition result data including target category, position and confidence level.

[0012] Preferably, the S5 comprises the following steps: S51, constructing a reinforcement learning decision model, inputting the target recognition result data and historical control records; S52, calculating the optimal motion path and action sequence, and generating preliminary control instructions; S53. Evaluate instruction safety based on the collision prediction algorithm. The calculation formula is: , in, is the safety risk factor, is the robot velocity vector, is the obstacle velocity vector, is the relative position vector; Safety threshold: ,exceeding the limit triggers re-planning; S54: Outputting the optimized robot motion control instructions.

[0013] Preferably, the S7 comprises the following steps: S71, collecting environmental feedback images and robot joint sensor data after the action is executed; S72, calculating the error matrix between the actual action and the expected target; S73, using a gradient descent algorithm to update the weight parameters of the reinforcement learning decision model; S74: Generate the adaptive control optimization data including new weights.

[0014] Preferably, the system includes a visual data acquisition module, a feature processing and analysis module, and a motion control execution module.

[0015] Preferably, the visual data acquisition module includes a multispectral camera unit, an inertial sensing unit and a data fusion unit; The multispectral camera unit collects ambient light and thermal map data; The inertial sensing unit monitors the robot's six-degree-of-freedom motion parameters in real time; The data fusion unit aligns spatiotemporal data and outputs a synchronized data set.

[0016] Preferably, the feature processing and analysis module includes an image preprocessing unit, a feature extraction unit and a recognition decision unit; The image pre-processing unit performs denoising, enhancement and geometric correction operations; The feature extraction unit deploys a lightweight neural network to achieve real-time feature encoding; The recognition decision unit integrates target detection and path planning algorithms to generate a control strategy.

[0017] Compared with the prior art, the present invention has the following beneficial effects: 1. In this invention, by providing a dynamic environment perception module, a real-time identification mechanism for ambient light interference is established during real-time grasping operations guided by robot vision. Differentiated compensation strategies are set for different working scenarios to ensure the stability of visual data acquisition under complex lighting conditions. Simultaneously, the compensated light parameters are synchronized to the image processing unit, which can reduce feature recognition distortion caused by abnormal environmental reflections in real time, ensure the accuracy of the robot's target positioning, and reduce trajectory errors guided by vision. 2. In the present invention, by providing a posture offset compensation module, when controlling the motion trajectory of the robot end effector, the real-time relative posture offset between the target object and the actuator is calculated, and the motion trajectory deviation is dynamically determined during the grasping process. This allows the system to avoid grasping failures caused by object displacement. When the posture offset is detected to exceed the safety threshold, the system can immediately generate posture correction instructions based on the preset kinematic model, allowing the robot actuator to calibrate the operation path in real time, ensuring the completion of high-precision assembly tasks. 3. In the present invention, by setting up a multi-objective decision optimization module, when performing the robot's multi-object collaborative grasping task, the three-dimensional spatial topological relationship of objects in the scene is automatically identified, the grasping priority weights of different objects are dynamically assigned, and the collision-free grasping sequence is planned in real time according to the weight allocation results, so that the system can avoid recognition confusion and operation conflicts between similar objects, reduce the collision risk caused by decision-making errors during multi-objective operations, and improve the safety of the robot operation system and the task success rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a flow chart of a robot vision system construction and control method of the present invention; Figure 2 This is a schematic diagram of a robot vision system and its control method according to the present invention. DETAILED DESCRIPTION

[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0020] See also Figure 1 The robot vision system is constructed and controlled by the method, and the method comprises the following steps: S1, collect environmental image data and robot motion state data; S2, preprocessing the environmental image data in real time to generate standardized image data; S3, extracting visual feature parameters based on the standardized image data and generating a visual feature descriptor; S4, performing target object recognition and positioning processing based on the visual feature descriptor and the robot motion state data to generate target recognition result data; S5. Generate robot motion control instructions based on target recognition result data and preset control strategy; S6, execute the robot motion control instructions and drive the robot actuator to complete the action; S7, dynamically adjust control strategy parameters based on action execution feedback data to generate adaptive control optimization data; S1 includes the following steps: S11, collecting visible light and infrared image data in the environment through a multispectral camera array; S12, collecting acceleration, angular velocity and attitude data of the robot through an inertial measurement unit; S13, fusing the visible light and infrared image data with the posture data to generate a spatiotemporally synchronized environmental image dataset; S14, performing timestamp alignment processing on the environmental image dataset to ensure consistency of data acquisition frequency; S2 includes the following steps: S21, performing denoising filtering on the environmental image data, and using an adaptive median filter to eliminate impulse noise; S22. Perform contrast enhancement processing on the denoised image data by using a limited contrast adaptive histogram equalization algorithm using the following formula: , in, is the normalized contrast, is the original image pixel value, is the average pixel value of the region, is the regional standard deviation; threshold value: time triggered additional enhancement mechanism; S23, geometric distortion correction is performed on the enhanced image data, and perspective transformation is performed based on the camera intrinsic parameter model; S24, normalize the corrected image data to a unified resolution to generate standardized image data; S3 includes the following steps: S31, using a convolutional neural network model to extract the spatial features of the standardized image data to generate a primary feature map; S32, generate feature descriptors through multi-scale pooling and perform feature fusion: , wherein, is the normalized fusion feature vector, is the spatial feature vector, is the frequency domain feature vector, is the weighting coefficient, default threshold ; S33, fuse the multi-resolution feature descriptors and the robot motion state data to construct a spatio-temporal joint feature vector and calculate the feature importance weight: , wherein, is the feature importance weight, is the norm of the current feature vector, is the maximum norm of all feature vectors; threshold value: ignore the feature; S34, reduce the dimension of the spatio-temporal joint feature vector through principal component analysis to generate a visual feature descriptor; S4 includes the following steps: S41, input the visual feature descriptor into the YOLOv7 target detection model to identify dynamic and static targets in the environment; S42, calculate the 3D position coordinates of the target object based on the principle of binocular vision triangulation; S43, combine the robot pose data to convert the target position to the robot coordinate system; S44, generate target recognition result data containing target category, position and confidence, and generate target similarity score: , wherein, is the target similarity score, ranging from [-1, 1], is the target object feature vector, is the reference template feature vector; Threshold: When it is determined to be a match; S5 includes the following steps: S51. Build a reinforcement learning decision model and input target recognition result data and historical control records; S52, calculating the optimal motion path and action sequence, and generating preliminary control instructions; S53. Evaluate instruction safety based on the collision prediction algorithm. The calculation formula is: , in, is the safety risk factor, is the robot velocity vector, is the obstacle velocity vector, is the relative position vector; Safety threshold: ,exceeding the limit triggers re-planning; S54, output optimized robot motion control instructions and execute path optimization: , in, To optimize the path index, length is the physical length of the path, time is the estimated execution time, and risk is the normalized risk value. is the weight coefficient; Threshold: , when risk>0.6 ; S7 includes the following steps: S71, collecting environmental feedback images and robot joint sensor data after the action is executed; S72, calculating the error matrix between the actual action and the expected target; S73. Use the gradient descent algorithm to update the weight parameters of the reinforcement learning decision model, and use the weight update formula to adjust the model parameters: , in, is the weight change, To control the error, is the input feature vector; Threshold: , the maximum number of iterations is 100; S74, generating adaptive control optimization data including new weights; The system includes a visual data acquisition module, a feature processing and analysis module, and a motion control execution module; The visual data acquisition module includes a multispectral camera unit, an inertial sensing unit, and a data fusion unit; The multispectral camera unit collects ambient light and thermal map data; The inertial sensing unit monitors the robot's six-degree-of-freedom motion parameters in real time; The data fusion unit aligns spatiotemporal data and outputs synchronized datasets; The feature processing and analysis module includes an image preprocessing unit, a feature extraction unit, and a recognition decision unit; The image preprocessing unit performs denoising, enhancement and geometric correction operations; The feature extraction unit deploys a lightweight neural network to achieve real-time feature encoding; The recognition decision unit integrates target detection and path planning algorithms to generate control strategies.

[0021] The steps of constructing a robot vision system and its control method are as follows: 1. Principle of environmental data collection: The system first collects environmental images and robot motion status data through a multispectral camera array and an inertial measurement unit. The camera captures visible light and infrared spectra to generate high-resolution image sequences, while the inertial unit monitors the robot's acceleration, angular velocity, and posture changes in real time. The data acquisition process emphasizes spatiotemporal synchronization to ensure that the image and motion data are strictly aligned in timestamps. This avoids the delay problems common in traditional systems and provides an original and consistent data foundation for subsequent processing.

[0022] 2. Image preprocessing principle: The original image data undergoes denoising filtering and contrast enhancement to reduce environmental interference. An adaptive median filter removes impulse noise, while a contrast-limited adaptive histogram equalization algorithm improves the recognizability of image details. The processed image undergoes geometric distortion correction and perspective transformation is adjusted based on the camera intrinsic parameter model to ensure that the image data is standardized to a uniform resolution. This step ensures the stability of the visual input and prevents subsequent feature extraction from deviating due to data quality issues.

[0023] 3. Principle of visual feature extraction: The standardized image is input into the convolutional neural network to extract spatial and frequency domain features. The network generates multi-scale feature descriptors through multi-layer convolution and pooling operations to capture the target's morphology, texture, and motion pattern. The feature fusion mechanism combines the robot's motion state data to construct a joint spatiotemporal feature vector. The principal component analysis method is used to reduce the dimension of the feature vector and generate a compact visual feature descriptor. This process simulates the hierarchical processing of biological vision, enhances the system's sensitivity to the key attributes of the target, and ensures the efficiency and discriminability of feature representation.

[0024] 4. Target recognition and positioning principles: The visual feature descriptor inputs the target detection model, identifies dynamic and static targets in the environment, the model calculates the target similarity score, compares the target features with the reference template through vector dot product, determines the target category and position, combines the principle of binocular vision triangulation, the system converts the target coordinates to the robot coordinate system, realizes 3D positioning, and introduces real-time feedback in the identification process. When the similarity score is lower than the threshold, re-detection is triggered to avoid misidentification, which ensures that the robot accurately perceives the target pose and provides reliable input for control decision.

[0025] 5. Control instruction generation principle: Based on the target recognition result and historical control record, the reinforcement learning decision model generates the preliminary motion path and action sequence, the model evaluates the path safety, calculates the relative risk of the robot speed vector and obstacles, and re-plans the path when the risk exceeds the limit. The optimization process balances the path length and execution time to generate robot motion control instructions. The instruction generation emphasizes real-time performance, the system dynamically adjusts the weight parameters to adapt to environmental changes, and ensures that the instructions achieve the optimal balance between efficiency and safety.

[0026] 6. Action execution principle: The robot execution mechanism receives control instructions to drive the robot to complete the specified action. The execution process monitors joint sensor data in real time to ensure that the action trajectory is consistent with the planning. When a pose deviation is detected, the system immediately corrects the path through the kinematics model to compensate for external disturbances. After action execution, the system collects feedback data to form a closed-loop control flow, which simulates the adaptive adjustment of human operation, improving the accuracy and robustness of task completion.

[0027] 7. Dynamic parameter optimization principle: The action execution feedback data dynamically adjusts the control strategy parameters, the system calculates the error matrix between the actual and expected values, updates the model weights through the gradient descent algorithm, and controls the learning rate within a low value range to avoid overfitting. At the same time, the number of iterations is limited to ensure efficient convergence of the optimization process. The generated adaptive control optimization data is fed back to the decision model in real time to form a continuous learning mechanism. This principle enables the system to gradually improve performance over time, adapt to new scenarios, and enhance overall intelligence and reliability.

[0028] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0029] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A robot vision system construction and control method thereof, characterized by: The method comprises the following steps: S1, collect environmental image data and robot motion state data; S2. Preprocess the environmental image data in real time to generate standardized image data; S3. Extracting visual feature parameters based on the standardized image data to generate a visual feature descriptor; S4, performing target object recognition and positioning processing according to the visual feature descriptor and the robot motion state data to generate target recognition result data; S5, generating robot motion control instructions based on the target recognition result data and a preset control strategy; S6, executing the robot motion control instruction to drive the robot actuator to complete the action; S7. Dynamically adjust control strategy parameters based on action execution feedback data to generate adaptive control optimization data.

2. A robot vision system construction and control method according to claim 1, characterized in that: Said S1 comprises the following steps: S11, collecting visible light and infrared image data in the environment through a multispectral camera array; S12, collecting acceleration, angular velocity and attitude data of the robot through an inertial measurement unit; S13, fusing the visible light and infrared image data with the posture data to generate a spatiotemporally synchronized environmental image data set; S14: performing timestamp alignment processing on the environmental image dataset to ensure consistency of data acquisition frequency.

3. The robot vision system construction and control method according to claim 1, characterized in that: The S2 comprises the following steps: S21, performing denoising filtering on the environmental image data, using an adaptive median filter to eliminate impulse noise; S22, performing contrast enhancement processing on the denoised image data, using a limited contrast adaptive histogram equalization algorithm; S23, performing geometric distortion correction on the enhanced image data, and performing perspective transformation based on the camera intrinsic parameter model; S24 , normalizing the corrected image data to a uniform resolution to generate the standardized image data.

4. The method for constructing and controlling a robot vision system according to claim 1, wherein: The S3 includes the following steps: S31, using a convolutional neural network model to extract spatial features of the standardized image data and generate a primary feature map; S32. Generate feature descriptors through multi-scale pooling and perform feature fusion: , in, is the normalized fusion feature vector, is the spatial eigenvector, is the frequency domain eigenvector, is the weighting coefficient, , the default threshold ; S33, fusing the multi-resolution feature descriptor with the robot motion state data to construct a spatiotemporal joint feature vector; S34. Reduce the dimension of the spatiotemporal joint feature vector by principal component analysis to generate the visual feature descriptor.

5. The robot vision system construction and control method according to claim 1, characterized in that: The S4 comprises the following steps: S41, inputting the visual feature descriptor into a YOLOv7 target detection model to identify dynamic and static targets in the environment; S42, calculating the 3D position coordinates of the target object based on the binocular vision triangulation principle; S43, combining the robot posture data, converting the target position to the robot coordinate system; S44: Generate target recognition result data including target category, position and confidence level.

6. The robot vision system construction and control method according to claim 1, characterized in that: The S5 comprises the following steps: S51, constructing a reinforcement learning decision model, inputting the target recognition result data and historical control records; S52, calculating the optimal motion path and action sequence, and generating preliminary control instructions; S53. Evaluate instruction safety based on the collision prediction algorithm. The calculation formula is: , in, is the safety risk factor, is the robot velocity vector, is the obstacle velocity vector, is the relative position vector; Safety threshold: ,exceeding the limit triggers re-planning; S54: Outputting the optimized robot motion control instructions.

7. The robot vision system construction and control method according to claim 1, characterized in that: The S7 comprises the following steps: S71, collecting environmental feedback images and robot joint sensor data after the action is executed; S72, calculating the error matrix between the actual action and the expected target; S73, using a gradient descent algorithm to update the weight parameters of the reinforcement learning decision model; S74: Generate the adaptive control optimization data including new weights.

8. A robot vision system, which realizes the robot vision system construction and control method according to claim 1, characterized in that: The system includes a visual data acquisition module, a feature processing and analysis module, and a motion control execution module.

9. The robot vision system according to claim 8, characterized in that: The visual data acquisition module includes a multispectral camera unit, an inertial sensing unit and a data fusion unit; The multispectral camera unit collects ambient light and thermal map data; The inertial sensing unit monitors the robot's six-degree-of-freedom motion parameters in real time; The data fusion unit aligns spatiotemporal data and outputs a synchronized data set.

10. The robot vision system according to claim 8, characterized in that: The feature processing and analysis module includes an image preprocessing unit, a feature extraction unit and a recognition decision unit; The image pre-processing unit performs denoising, enhancement and geometric correction operations; The feature extraction unit deploys a lightweight neural network to achieve real-time feature encoding; The recognition decision unit integrates target detection and path planning algorithms to generate a control strategy.

Citation Information

Cited By

  • Thickness detection device for panel display substrate

    CN121089599A

  • Robot visual servo control method for operating biochemical instrument

    CN122077657A