Panoramic perception method based on six-degree-of-freedom bionic eye and target focusing device
By employing a panoramic perception method based on a six-degree-of-freedom bionic eye, and combining information fusion from binocular vision sensors and lidar, the challenges of perception accuracy and recognition in dynamic environments in traditional binocular vision systems for industrial applications have been solved. This has enabled high-precision, all-around target detection and focusing, thereby improving production efficiency and product quality.
Patent Information
- Application Number
- CN202511089601.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-11-21
AI Technical Summary
Traditional binocular vision cameras suffer from insufficient stereo vision perception accuracy, weak dynamic scene perception capability, and difficulty in recognizing single texture areas in industrial applications. This results in insufficient dynamic tracking and real-time adjustment capabilities, as well as obstacles in recognizing complex surfaces and special materials, making it difficult to meet the needs of high-precision assembly and inspection.
A panoramic perception method based on a six-degree-of-freedom bionic eye is adopted. By fusing binocular vision sensors and LiDAR, and utilizing a logic block development kit, adaptive zoom algorithm and neck-eye collaborative control, multi-sensor information fusion and precise focusing are achieved. Combined with a deep learning target detection model, target features and motion state are processed.
Significantly enhances visual perception capabilities, enabling 360°×270° all-around perception, reducing depth measurement errors, improving target recognition and tracking capabilities in complex environments, enhancing the detection of special materials, optimizing focusing performance, improving production efficiency and product quality, reducing maintenance costs, and expanding application areas.
Smart Images

Figure CN120997467A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer vision, in particular to a panoramic perception method based on a six-degree-of-freedom bionic eye and a target focusing device. BACKGROUND
[0002] With the rapid development of intelligent manufacturing, machine vision technology, as the "eyes" of industrial automation, is rapidly popularizing worldwide.
[0003] At present, traditional binocular vision cameras face serious limitations in industrial applications, mainly in the following three aspects:
[0004] 1) Insufficient stereo vision perception accuracy: Taking mainstream products such as Intel RealSense D455 (the fourth generation of stereo vision depth camera in the Intel RealSense D400 series), Microsoft Kinect v2 (the second generation of 3D body sensing depth camera launched by Microsoft in 2014), etc. as examples, their performance in minimum depth, RGB frame rate, and visual range is relatively low. Test data shows that under standard lighting conditions (1000 lux), the ranging error of the traditional scheme is about 3.2 cm, and in a dark environment (10 lux) it even fails completely.
[0005] 2) Weak dynamic scene perception ability: In an environment with motion or vibration, the traditional vision system has difficulty in accurately positioning and tracking the target. Tests show that under the condition of 2Hz dynamic jitter, the ranging error is as high as 8.7cm, seriously affecting the accuracy of recognition and positioning.
[0006] 3) Difficulty in recognizing single-texture areas: For objects with single or repetitive surface texture, existing technologies have difficulty in effectively identifying and distinguishing them, resulting in "inaccurate, unable to follow, and weak recognition" in many industrial scenarios.
[0007] Two, specific pain points in practical applications
[0008] 1) Insufficient dynamic tracking and real-time adjustment capability, mainly in the following two aspects:
[0009] In high-precision assembly operations, vibrations in the industrial environment cause relative motion between the target object and the robot end effector, and the traditional vision system has insufficient response speed (usually > 50ms), which cannot compensate for this small displacement in real time, resulting in a decrease in assembly precision.
[0010] The lack of high-speed and high-precision three-dimensional perception capability makes it impossible for robots to make real-time adjustments based on the small changes in the actual position of the workpiece. According to industry tests, the conventional industrial robot has a positioning accuracy decay of up to 30% without high-precision vision feedback, which seriously affects the assembly yield of precision parts.
[0011] 2) Complex surface and special material recognition obstacles
[0012] In the process of precision manufacturing, the detection ability of traditional binocular vision system for glass, transparent plastic and other materials is very low, and the recognition rate is usually less than 40%, and such materials are very common in medical devices and optical element manufacturing.
[0013] Surface defect detection limitations: Micro surface defects (such as cracks or pits less than 50 microns) are difficult to detect under normal light. Industry data shows that in the detection of aerospace parts, the traditional visual system has a high false negative rate of 18% for micro cracks, which is far from meeting the safety standard requirements.
[0014] In view of the problems in the related art, no effective solution has been proposed so far. SUMMARY
[0015] In view of the problems in the related art, the present application proposes a panoramic perception method based on a six-degree-of-freedom bionic eye and a target focusing device to overcome the above technical problems existing in the prior art.
[0016] To this end, the specific technical solutions adopted by the present application are as follows:
[0017] In a first aspect, the present application provides a panoramic perception method based on a six-degree-of-freedom bionic eye, comprising:
[0018] S1, real-time acquisition of a synchronous image pair based on a binocular vision sensor, and calculation of target position information using a binocular perception algorithm;
[0019] S2, real-time scanning using a laser radar, pre-processing of target position cloud data, obtaining target point cloud coordinates based on a Euclidean clustering algorithm, and converting the coordinates into target position information of the binocular vision sensor;
[0020] S3, using a logic block development kit, fusing the local geometric shape features in the target position information obtained by the binocular vision sensor and the laser radar respectively to generate a shallow thermal map; using an adaptive zoom algorithm, fusing the global geometric shape features in the target position information obtained by the binocular vision sensor and the laser radar respectively to generate a deep thermal map;
[0021] S4, using an attention mechanism and a feature weighting algorithm to fuse the local features in the shallow thermal map and the deep thermal map to generate a feature matrix, and combining a lens imaging formula to calculate the focal length, and using a controller to adjust the target focal length of the binocular vision sensor in real time, and verifying the target focal length accuracy by means of a sharpness detection model;
[0022] S5, based on the target position information calculated by the binocular perception algorithm, the interpupillary distance is dynamically adjusted to match the target position information by using binocular stereo vision. The target interpupillary distance of the binocular vision sensor is adjusted in real time through the controller, and the interpupillary distance feature is enhanced through the unsupervised reconstruction mechanism;
[0023] S6, based on the neck-eye collaborative control algorithm, a bionic motion model is constructed to extract target features and motion states, and a deep learning target detection model and a motion perception algorithm are used to process the target features and motion states respectively to optimize the target position information.
[0024] Further, based on the binocular vision sensor, real-time synchronous image pairs are obtained, the target position information is calculated by using the binocular perception algorithm, and an initial disparity map is constructed. The distance of the target position includes:
[0025] S11, using the binocular perception algorithm to construct an iterative geometric coding cost volume to fuse the geometric shape features of the synchronous image pairs;
[0026] S12, using the gated recurrent unit to perform multi-round optimization on the synchronous image pairs to calculate the distance of the target object to the binocular vision sensor.
[0027] Further, the target position cloud data is preprocessed by using the laser radar to scan in real time, the target point cloud coordinates are obtained based on the Euclidean clustering algorithm, and the target position information of the binocular vision sensor is converted by using the coordinate system including:
[0028] S21, using the laser radar to emit modulated laser beams and receive reflected signals, and calculating the target position original distance data by using the time of flight principle;
[0029] S22, removing noise points and downsampling the obtained target position cloud data by statistical outlier filtering and voxel grid filtering;
[0030] S23, based on the Euclidean clustering algorithm, the target point cloud coordinates are segmented, the target point cloud clusters are extracted, and the three-dimensional coordinates of the cluster center points are calculated;
[0031] S24, converting the three-dimensional coordinates of the laser radar into the target position information of the binocular vision sensor through the coordinate system.
[0032] Further, using the logic block development kit, the local geometric shape features in the target position information obtained by the binocular vision sensor and the laser radar are fused to generate a shallow heat map; using the adaptive zoom algorithm, the global geometric shape features in the target position information obtained by the binocular vision sensor and the laser radar are fused to generate a deep heat map including:
[0033] S31. Utilize deep learning multi-scale feature fusion to fuse the target position information acquired by the LiDAR with the target position information acquired by the binocular vision sensor.
[0034] S32. Based on the local and global geometric features in the target location information, distinguish between shallow and deep heat maps.
[0035] Furthermore, by fusing local features from shallow and deep heatmaps using an attention mechanism and feature weighting algorithm, a feature matrix is generated. The focal length is then calculated using a lens imaging formula. The target focal length of the binocular vision sensor is adjusted in real-time by a controller. The accuracy of the target focal length is verified using a sharpness detection model, including:
[0036] S41. Utilize the attention mechanism to autonomously learn local features in shallow and deep heatmaps, and use a feature weighting algorithm to weight and fuse the extracted local features to generate a matrix containing target location depth and spatial location information.
[0037] S42. Based on the weighted fusion matrix, the focal length is calculated in real time using the lens imaging formula, and the calculated focal length is converted into a control signal for the controller.
[0038] S43. The controller adjusts the target focal length position of the binocular vision sensor in real time according to the control signal.
[0039] S44. By using a sharpness detection model to detect the sharpness of the target position information in real time, the accuracy of target focus adjustment can be further optimized.
[0040] Furthermore, based on the target position information calculated by the binocular perception algorithm, the interpupillary distance is dynamically adjusted to match the target position information using binocular stereo vision. The target interpupillary distance of the binocular vision sensor is adjusted in real time by the controller, and the interpupillary distance features are enhanced through an unsupervised reconstruction mechanism, including:
[0041] S51. Generate an initial disparity map based on the binocular perception algorithm, extract target position information, combine the geometric constraints of binocular stereo vision, dynamically adjust the interpupillary distance to match the target position information, and adjust the target interpupillary distance of the binocular vision sensor in real time through the controller.
[0042] S52. Utilize an unsupervised reconstruction mechanism to construct a multi-scale feature pyramid and enhance the features of weakly textured regions through adaptive enhancement of the cyclic neighborhood.
[0043] Furthermore, based on the neck-eye coordinated control algorithm, a biomimetic motion model is constructed to extract target features and motion states. A deep learning target detection model and a motion perception algorithm are then used to process the target features and motion states respectively, optimizing the target position information, including:
[0044] S61, based on the neck-eye cooperative control algorithm, a bionic motion model containing binocular vision information and neck motion information is constructed, and target features and motion states are extracted from the binocular vision information and the neck motion information;
[0045] S62, a deep learning target detection model is used to detect moving target features, to obtain detection frame coordinates and feature vectors, and to input them to a single target tracking module for processing to generate a continuous target feature sequence;
[0046] S63, a motion perception algorithm is used to calculate the motion state parameters of the target in real time, to obtain the calculation results, and to input them to a space-time behavior analysis unit for hierarchical response processing.
[0047] Further, a deep learning target detection model is used to detect moving target features, to obtain detection frame coordinates and feature vectors, and to input them to a single target tracking module for processing to generate a continuous target feature sequence, which includes:
[0048] S621, the detection frame coordinates and the feature vectors are input to a single target tracking module, the target trajectory is predicted by a Kalman filter, and the Hungarian algorithm is used for trajectory association to generate a continuous target feature sequence.
[0049] Further, a motion perception algorithm is used to calculate the motion state parameters of the target in real time, to obtain the calculation results, and to input them to a space-time behavior analysis unit for hierarchical response processing, which includes:
[0050] S631, the calculation results are input to a space-time behavior analysis unit, a hidden Markov model is used to identify target behavior patterns, and abnormal behavior events will be triggered to a warning decision system for hierarchical response processing.
[0051] In a second aspect, the application also provides a target focusing device based on a six-degree-of-freedom bionic eye, which includes:
[0052] a base;
[0053] a neck yaw degree of freedom arranged at the top end of the base;
[0054] a neck connecting piece arranged at the top of the neck yaw degree of freedom
[0055] a neck-eye connecting piece movably connected to the top of the neck connecting piece;
[0056] an eye yaw degree of freedom, an eye yaw connecting piece, an eye pitch connecting piece and eyeballs, which are symmetrically arranged in order from bottom to top at the top of the neck-eye connecting piece;
[0057] a laser radar arranged at the top end of the neck-eye connecting piece and located between the two groups of eyeballs.
[0058] The application has the following advantages:
[0059] I. Technical Effects
[0060] 1) Significant improvement in visual perception ability: Traditional machine vision systems have limited field of view angles, such as Intel RealSense D455, which only has a 87° x 58° viewing angle range, making it difficult to provide all-around perception. This invention simulates the human eye and neck coordination mechanism, achieving 360° x 270° all-around perception, which is a significant expansion of the traditional system's visual range. In practical application scenarios, such as spacecraft integrated circuit defect detection, traditional vision systems may miss defects in the corners or edges due to limited viewing angles, while this invention can detect the entire integrated circuit without dead angles, significantly improving the comprehensiveness of detection.
[0061] In terms of depth perception accuracy, the measurement error of traditional single sensor depth is usually above ±0.5mm, with a ranging error of about 3.2cm under standard lighting and an error of up to 8.7cm in a dynamic jitter (2Hz) environment. This invention significantly reduces the depth measurement error under standard lighting through multi-sensor fusion technology, and the error control effect is more significant in a dynamic jitter environment. This enables more accurate acquisition of three-dimensional information of objects in precision manufacturing and other fields, such as real-time adjustment based on the slight changes in the actual position of workpieces in precision part assembly, avoiding the problem of reduced assembly precision caused by inaccurate depth perception, and effectively improving assembly quality.
[0062] 2) Significant improvement in adaptability to complex environments: This invention can achieve stable and accurate focusing under complex conditions such as alternating near and far targets through laser radar guided focusing and adaptive zooming algorithms, significantly improving the detection accuracy of small targets. At the same time, in dynamic scenarios, through the development of neck-eye coordination control algorithms based on bionic mechanisms, accurate target recognition and tracking are ensured in dynamic scenarios, significantly improving dynamic target recognition rate and significantly reducing perception error. For example, in high-altitude target detection for unmanned aerial vehicles, even in dim light and high-speed target movement, this invention can still accurately lock and track the target with high real-time tracking frame rate and excellent dynamic target recognition rate, while traditional systems are difficult to meet the operational requirements in such complex environments.
[0063] 3) Breakthrough in recognition of special materials and single texture: Traditional vision systems have very low recognition ability for objects with single or repetitive surface texture, as well as transparent and highly reflective materials. This invention realizes multi-modal data fusion of laser radar and high-definition cameras, and through the fusion of point cloud and image information, it obtains more rich texture information, effectively solving the problem of single texture area recognition. In the process of precision manufacturing, for materials such as glass and transparent plastic commonly used in medical devices and optical component manufacturing, this invention can significantly improve detection capability.
[0064] 4) Focus performance optimization: Traditional vision systems have limited focus depth range and lack the mechanism to achieve precise focusing in complex environments. The focusing adjustment process is slow. The invention achieves precise focusing in complex environments, significantly shortens the focus response time in intelligent chemical experiment systems, greatly improves work efficiency, and can simultaneously maintain clear imaging at different height levels, meeting the needs of high-speed precision operations.
[0065] II. Economic effect
[0066] 1) Improve production efficiency and product quality: In the field of industrial manufacturing, the invention can effectively solve the problem of insufficient dynamic tracking and real-time adjustment capability of traditional vision systems. In high-precision assembly operations, the response speed of traditional vision systems is insufficient (usually > 50 ms), leading to a decline in assembly precision. Conventional industrial robots without high-precision vision feedback have a positioning accuracy decay of up to 30%. However, the invention, with its fast response speed and precise vision perception, can enable robots to adjust in real time based on the actual position of the workpiece, reducing assembly errors and improving assembly yield. For example, in electronic component assembly, it can reduce the scrap rate, reduce material waste and rework costs, thereby improving production efficiency and bringing direct economic benefits to enterprises.
[0067] 2) Reduce maintenance costs: Due to its higher reliability and stability, the invention has a lower failure rate in complex environments compared to traditional vision systems. In scenarios such as aerospace part detection, where equipment stability is highly required, traditional vision systems often need frequent maintenance and calibration due to insufficient detection capability for minor surface defects and performance degradation in dynamic environments, increasing operational costs for enterprises. The invention reduces equipment failure and maintenance frequency, reducing enterprise investment in equipment maintenance.
[0068] 3) Expand application fields and market space: The high performance of the invention enables it to meet the needs of special operation scenarios, such as spacecraft integrated circuit defect detection, intelligent chemical experiment systems, and unmanned aerial vehicle high-altitude target detection. These new application fields open up new market opportunities for enterprises, which can obtain more business orders, expand market share, and increase economic benefits by providing products and services based on the invention technology. At the same time, as the market demand for high-precision and high-adaptability machine vision systems continues to grow, the invention is expected to drive the development of related industries and form new economic growth points.
[0069] III. Social effect
[0070] 1) Promote the upgrading of intelligent manufacturing industry: the appearance of the invention provides advanced visual perception technology for the field of intelligent manufacturing, which can effectively overcome the limitations of existing technology and significantly improve the visual perception ability of robots in fine operation. It helps to promote the development of manufacturing industry towards intelligent and automated direction, improve the overall competitiveness of China's manufacturing industry, and promote the optimization and upgrading of industrial structure.
[0071] 2) Ensure product quality and safety: in the field of aerospace, medical devices and other high-quality and safety requirements, the invention can improve detection accuracy and reliability. In the detection of aerospace parts, the miss detection rate of micro cracks is reduced, effectively ensuring product quality and safety, and reducing safety accidents caused by product quality problems.
[0072] 3) Promote green production and sustainable development: by improving production efficiency and product quality, reducing scrap rate and rework rate, the invention helps to reduce resource consumption and environmental pollution, in line with the concept of green production and sustainable development. In today's increasingly scarce resources and increasingly serious environmental problems, promoting enterprises to adopt more efficient and environmentally friendly production technology has a positive role in realizing the coordinated development of economy and environment. BRIEF DESCRIPTION OF DRAWINGS
[0073] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0074] Figure 1 is a flow chart of a panoramic perception method based on a six-degree-of-freedom bionic eye according to an embodiment of the present application;
[0075] Figure 2 is one of the structural schematic diagrams of a target focusing device based on a six-degree-of-freedom bionic eye according to an embodiment of the present application;
[0076] Figure 3 is a front view of a target focusing device based on a six-degree-of-freedom bionic eye according to an embodiment of the present application.
[0077] Figure 4 is a front view of a target focusing device based on a six-degree-of-freedom bionic eye according to an embodiment of the present application.
[0078] In the drawings:
[0079] 1, base; 2, neck yaw freedom; 3, neck connector; 4, eye connector; 5, laser radar; 6, eye yaw freedom; 7, eye yaw connector; 8, eye pitch connector; 9, eyeball. DETAILED DESCRIPTION
[0080] To further illustrate the embodiments, the present application provides drawings, which are part of the disclosure of the present application, mainly used to illustrate the embodiments, and can be explained in conjunction with the related description of the specification to understand the operating principle of the embodiments. With reference to these contents, those skilled in the art should understand other possible implementations and advantages of the present application.
[0081] According to an embodiment of the present application, a panoramic perception method and target focusing device based on a six-degree-of-freedom bionic eye are provided.
[0082] The present application will be further described in conjunction with the drawings and specific embodiments, as shown in Figure 1 A panoramic perception method based on a six-degree-of-freedom bionic eye according to an embodiment of the present application includes:
[0083] S1, real-time acquisition of a synchronous image pair based on a binocular vision sensor (i.e., left and right high-definition CMOS sensors, two eyeballs 9), and calculation of target position information using a binocular perception algorithm (i.e., IGEV (Iterative Geometry Encoding Volume));
[0084] S2, real-time scanning of target position cloud data using a laser radar 5 for preprocessing, acquisition of target point cloud coordinates based on a Euclidean clustering algorithm, and conversion of the coordinates into target position information of the binocular vision sensor;
[0085] S3, fusion of local geometric shape features in the target position information obtained by the binocular vision sensor and the laser radar 5 respectively using a logic block development kit (i.e., CLB development kit) to generate a shallow thermal map (i.e., shallow feature fusion); fusion of global geometric shape features in the target position information obtained by the binocular vision sensor and the laser radar 5 respectively using an adaptive zoom algorithm (i.e., Point-Vision) to generate a deep thermal map (i.e., deep feature fusion);
[0086] S4, fusion of local features in the shallow thermal map and the deep thermal map using an attention mechanism and a feature weighting algorithm to generate a feature matrix, and calculation of focal length in combination with a lens imaging formula, real-time adjustment of the target focal length of the binocular vision sensor through a controller (i.e., a PID controller), and verification of the target focal length accuracy through a sharpness detection model (i.e., Jetson Orin Nano);
[0087] S5, based on the target position information calculated by the binocular perception algorithm, the interpupillary distance is dynamically adjusted to match the target position information by using binocular stereo vision. The target interpupillary distance of the binocular vision sensor is adjusted in real time through the controller, and the interpupillary distance feature is enhanced through the unsupervised reconstruction mechanism;
[0088] S6, based on the neck-eye collaborative control algorithm, a bionic motion model is constructed to extract target features and motion states, and a deep learning target detection model (i.e., YOLOv9) and a motion perception algorithm (i.e., CVMS algorithm) are used to process target features and motion states respectively to optimize target position information.
[0089] In this alternative embodiment, based on the binocular vision sensor, real-time acquisition of synchronous image pairs is used to calculate target position information by using a binocular perception algorithm, and an initial disparity map is constructed to extract the distance of the target position, including:
[0090] S11, using the binocular perception algorithm to construct an iterative geometry encoding cost volume to fuse the geometric shape features of the synchronous image pairs;
[0091] S12, using a gated recurrent unit to perform multi-round optimization on the synchronous image pairs to calculate the distance of the target object to the binocular vision sensor.
[0092] It should be noted that the IGEV (Iterative Geometry Encoding Volume) binocular perception algorithm is used to accurately calculate the distance of the target object to the bionic eye system (target position information) based on the synchronous image pairs captured by the left and right high-definition CMOS sensors (binocular vision sensor). The IGEV algorithm constructs an iterative geometry encoding cost volume (Iterative Geometry Encoding Volume) to effectively fuse image features and geometric constraints, and uses a gated recurrent unit (GRU) for multi-round optimization, significantly improving the robustness and accuracy of depth estimation in weak texture areas and areas with discontinuous disparity. When deployed in a new environment, a small amount of samples can be used to quickly optimize model parameters and improve the accuracy of depth estimation. The output is a high-precision depth map of the target area and the corresponding confidence information, which provides key distance basis for the subsequent focusing step.
[0093] In this alternative embodiment, the target position cloud data is preprocessed by using the laser radar 5 to scan in real time, the target point cloud coordinates are obtained based on the Euclidean clustering algorithm, and the coordinates are converted into target position information of the binocular vision sensor, including:
[0094] S21, using the laser radar 5 to emit modulated laser beams and receive reflected signals, and calculating the original distance data of the target position by the time-of-flight principle;
[0095] S22, remove noise points and downsample the acquired target position cloud data by statistical outlier filtering and voxel grid filtering;
[0096] S23, segment the target point cloud coordinates based on the Euclidean clustering algorithm, extract the target point cloud cluster, and calculate the three-dimensional coordinates of the cluster center point;
[0097] S24, convert the three-dimensional coordinates of the laser radar 5 into target position information of the binocular vision sensor through the coordinate system.
[0098] It should be noted that the laser radar 5 emits a modulated laser beam and receives a reflected signal, and calculates the original distance data by the time of flight (ToF) principle. The original point cloud data is preprocessed: 1) using statistical outlier filtering (Statistical Outlier Removal, SOR) to remove noise points; 2) downsample by voxel grid filtering (VoxelGridFilter) to balance data density and computing efficiency; 3) based on the Euclidean clustering algorithm (DBSCAN) to segment the point cloud, extract the target object point cloud cluster, and calculate the three-dimensional coordinates (x, y, z) of the cluster center point; 4) through coordinate system conversion (laser radar 5 coordinate system→camera coordinate system), convert the target point cloud coordinates into depth values in the camera reference system, and output as a depth matrix.
[0099] The laser radar 5 measures the distance between it and the target, while the camera captures the image. The distance measured by the laser radar 5 through pulse ranging method is judged, if the distance is less than 1.2m, then according to the estimated depth of the last step, only the camera is used for adaptive focusing; if the distance is greater than or equal to 1.2m, the distance measured by the laser radar 5 is used to guide the camera focusing, which needs to use coordinate system conversion to convert the distance information measured by the laser radar 5 into the distance between the target and the camera, and then let the camera focus.
[0100] In this optional embodiment, the local geometric shape features in the target position information obtained by the binocular vision sensor and the laser radar 5 are fused to generate a shallow heat map using a logic block development kit; the global geometric shape features in the target position information obtained by the binocular vision sensor and the laser radar 5 are fused to generate a deep heat map using an adaptive zoom algorithm, including:
[0101] S31, fuse the target position information obtained by the laser radar 5 and the target position information obtained by the binocular vision sensor by using a deep learning multi-scale feature fusion method;
[0102] S32, according to the local geometric shape features and the global geometric shape features in the target position information, distinguish the shallow heat map and the deep heat map.
[0103] It should be noted that after adjusting the focal length, the point cloud information collected by the laser radar 5 and the image information collected by the camera are fused in a deep learning multi-scale feature fusion manner. In the shallow feature fusion, the edge features of the image and the local geometric shape features of the point cloud are fused; in the deep feature fusion, the high-level semantic features of the image and the global shape features of the point cloud are fused to enhance the recognition and understanding of the object.
[0104] When the detected target distance is <1.2m, the Point-Vision algorithm is triggered: the point cloud data and the image features are spatially aligned through the CLB development kit (logic block development kit) to generate a fused depth heat map. The "fusion result" refers to the fused "depth heat map" generated by the Point-Vision adaptive zoom algorithm.
[0105] In this optional embodiment, the local features in the shallow heat map and the deep heat map are fused by using an attention mechanism and a feature weighting algorithm to generate a feature matrix, and the focal length is calculated in combination with the lens imaging formula. The target focal length of the binocular vision sensor is adjusted in real time by the controller, and the target focal length accuracy is verified by a sharpness detection model, including:
[0106] S41, the local features in the shallow heat map and the deep heat map are autonomously learned by using an attention mechanism, and the extracted local features are weighted and fused by a feature weighting algorithm to generate a matrix containing target position depth and spatial position information;
[0107] S42, according to the matrix after weighted fusion, the focal length is calculated in real time in combination with the lens imaging formula, and the real-time calculated focal length is converted into a control signal of the controller;
[0108] S43, the target focal length position of the binocular vision sensor is adjusted in real time according to the control signal control of the controller;
[0109] S44, the sharpness of the target position information is detected in real time by means of a sharpness detection model, and the accuracy of the target focal length adjustment is further optimized.
[0110] It should be noted that the heat map combines the point cloud data of the laser radar 5 and the image information collected by the camera, and through feature fusion and attention mechanism weighting, a matrix containing accurate depth and spatial position information of the target object is finally generated. This result provides a key basis for subsequent focus adjustment. According to the fusion result, the focal length is adjusted by driving the 2-DOF eyeball motor according to the fused depth heat map, and the system adjusts the focal length by the following steps:
[0111] 1) Depth information extraction: the center point depth value (denoted as d) of the target object is extracted from the depth heat map.
[0112] 2) Focus calculation: According to the lens imaging formula 1 / f = 1 / u + 1 / v, where u is the object distance (i.e. depth value d) and v is the image distance, the required focal length f is calculated in combination with the camera parameters.
[0113] 3) Motor control: Convert the calculated focal length ff into a control signal for the eye motor (such as pulse number or angle), and use a PID controller to adjust the motor rotation angle in real time to move the camera lens to the target focal length position.
[0114] 4) Closed-loop verification: Real-time detection of image sharpness (such as gradient-based sharpness evaluation function) by Jetson Orin Nano forms a closed-loop feedback to further optimize the accuracy of focal length adjustment.
[0115] Using attention mechanism to automatically learn the most important part of the two modal features, and then weighting and fusing these important feature parts. The specific method is: splice the extracted image features and point cloud features, and then perform linear transformation through a fully connected layer to obtain a fused feature vector F. Through a multi-layer perceptron, F is nonlinearly transformed to output an attention weight matrix A with the same dimension as the original feature, whose element value is between 0 and 1, indicating the importance of the feature at the corresponding position or channel. Then normalize A to make the sum of its elements equal to 1.
[0116] Effect: Actual measurement shows that the focusing error is less than 0.03mm at a distance of 0.5m, which is 8 times higher than Intel RealSense D455 (Intel high-precision stereo vision depth camera).
[0117] In this optional embodiment, based on the target position information calculated by the binocular perception algorithm, the interpupillary distance is dynamically adjusted to match the target position information, the target interpupillary distance of the binocular vision sensor is adjusted in real time through the controller, and the interpupillary distance features are enhanced through the unsupervised reconstruction mechanism, including:
[0118] S51, generate an initial disparity map based on a binocular perception algorithm, extract target position information, dynamically adjust the interpupillary distance to match the target position information based on the geometric constraints of binocular stereo vision, and adjust the target interpupillary distance of the binocular vision sensor in real time through the controller;
[0119] S52, use an unsupervised reconstruction mechanism to construct a multi-scale feature pyramid, and enhance the features of weak texture regions through cyclic domain adaptation.
[0120] It should be noted that a high-precision depth map is generated by the IGEV algorithm to extract the distance dd of the target object. According to the geometric constraints of binocular stereo vision, the interpupillary distance BB is dynamically adjusted to match the target distance dd, ensuring that the parallax Δ is within the optimal range (usually Δmin≤Δ≤Δmax), and the calculation formula is:
[0121] B=k·d;
[0122] In the formula, k represents the proportionality coefficient (the empirical value is usually 0.05-0.1), which ensures that the parallax Δ=B·f / d is within a reasonable range; f represents the focal length.
[0123] The horizontal movement of the eyeball is driven by a servo motor (such as MAXON EC60), and the actual interpupillary distance is adjusted to the calculated value B. The control signal is based on the PID algorithm, which compensates for mechanical errors in real time.
[0124] Dynamic range limit: The interpupillary distance adjustment needs to be within the range allowed by the hardware (such as 60mm≤B≤80mm), and when it exceeds, it is compensated by the neck freedom degree.
[0125] Multi-modal collaboration: In weak texture scenes, combine the depth data of laser radar 5 to correct the interpupillary distance, and avoid the estimation deviation caused by insufficient image features.
[0126] Parallax optimization objective function: minimize∣Δ-Δideal∣, where Δideal=B·f / d.
[0127] Output three-dimensional point cloud data with confidence.
[0128] In this optional embodiment, based on the neck-eye collaborative control algorithm, a bionic motion model is constructed to extract target features and motion states, and a deep learning target detection model and a motion perception algorithm are used to process target features and motion states respectively, and the optimized target position information includes:
[0129] S61, based on the neck-eye collaborative control algorithm, a bionic motion model containing binocular vision information and neck motion information is constructed, and target features and motion states are extracted from the binocular vision information and neck motion information;
[0130] S62, a deep learning target detection model is used to detect moving target features to obtain detection box coordinates and feature vectors, and input them into a single target tracking module for processing to generate a continuous target feature sequence;
[0131] S63, a motion perception algorithm is used to calculate the motion state parameters of the target in real time to obtain the calculation results, and input them into a space-time behavior analysis unit for hierarchical response processing.
[0132] It should be noted that the YOLOv9 deep learning model (a deep learning target detection model) is deployed on the Jetson Orin AI processor to perform high-speed and high-precision target recognition on real-time image streams collected by left and right high-definition CMOS sensors. The trained YOLOv9 model is optimized for lightweight (channel pruning, 8-bit quantization) to meet the real-time requirements (≥30fps) of the embedded platform. The quantized YOLOv9 model outputs the precise bounding box coordinates, class labels, and confidence scores of the target object. This result serves as the basic input for subsequent depth estimation, focusing, and tracking processes. In particular, the bounding box coordinates are used to accurately locate the target region, effectively improving the efficiency and accuracy of subsequent depth calculation and focusing, and providing target position information for dynamic tracking.
[0133] According to the target, a bionic mechanism of neck-eye collaborative control algorithm is proposed, which realizes accurate target recognition and tracking in dynamic scenes through real-time calculation of target position and dynamic adjustment of motor control parameters. The specific method is as follows:
[0134] 1) Bionic motion model establishment: A bionic motion model containing binocular vision information and neck movement information is constructed to provide a basic framework for subsequent target recognition and tracking.
[0135] 2) Target feature extraction and motion state estimation: Extract target features from binocular vision information and estimate the motion state of the target by combining neck movement information to determine the position, speed, direction, and other key information of the target.
[0136] 3) Position coordinate calculation: According to the target features and motion state, the position coordinates of the target in space are accurately calculated to provide accurate target position information for motor control. The specific calculation formula is as follows: Based on the triangulation principle of stereo vision, the focal length of the left and right cameras is f, the baseline distance is b, and the left and right image point coordinates are (xl, y) and (xr, y) respectively. The target depth distance Z = f·b / (xl-xr), and the target position in the camera coordinate system is X = (xl+xr) / 2·Z / f; Y = y·Z / f.
[0137] In the formula, f represents the focal length, which refers to the focal length of the left and right cameras, and the unit is usually pixel; b represents the baseline distance, which refers to the distance between the optical centers of the left and right two cameras, and the unit is usually millimeter or meter, which is consistent with the unit of the focal length; xl represents the horizontal component of the left image point coordinate, which is the horizontal pixel coordinate of the target on the left camera image sensor; xr represents the horizontal component of the right image point coordinate, which is the horizontal pixel coordinate of the target on the right camera image sensor; y represents the vertical component of the image point coordinate, which is the vertical pixel coordinate of the target on the left and right camera image sensors (assuming that the left and right cameras are aligned in the vertical direction, so the vertical coordinates are the same); Z represents the target depth distance (Depth), which is the distance from the target to the camera plane, that is, the coordinate of the target on the Z axis of the camera coordinate system, and the unit is the same as the baseline distance; X represents the horizontal coordinate of the target in the camera coordinate system, which is the coordinate of the target on the X axis of the camera coordinate system, and the unit is the same as the baseline distance; Y represents the vertical coordinate of the target in the camera coordinate system, which is the coordinate of the target on the Y axis of the camera coordinate system, and the unit is the same as the baseline distance.
[0138] 4) Precise target tracking: using servo motor control equation and cooperative control rate, precise control of the motor is realized, and the cooperative movement of the eyeball and the neck is realized, and the target is continuously and stably tracked. The control equation of the servo motor adopts the PID control algorithm, the target angle is θd, the actual angle is θa, the error e = θd- θa, and the control voltage is:
[0139]
[0140] In the formula, u represents the control voltage (unit: volt V); K p represents the proportional gain coefficient (dimensionless); e represents the angle error (unit: radian rad), e = θd- θa; K i represents the integral gain coefficient (unit: s -1 ); ∫edt represents the time integral of the error (unit: rad·s); K d represents the differential gain coefficient (unit: s); de / dt represents the error change rate (unit: rad / s).
[0141] In this optional embodiment, the moving target features are detected by using a deep learning target detection model to obtain detection box coordinates and feature vectors, and input to a single target tracking module for processing to generate a continuous target feature sequence, including:
[0142] S621, input the detection box coordinates and feature vectors to the single target tracking module, predict the target trajectory through the Kalman filter, and use the Hungarian algorithm for trajectory association to generate a continuous target feature sequence.
[0143] It should be noted that the detection frame coordinates and the feature vector of the moving target detected by the YOLOV9 are transmitted to the single target tracking module for processing. The module predicts the target trajectory through the Kalman filter and uses the Hungarian algorithm for trajectory association to generate a continuous target ID sequence (target feature sequence).
[0144] In this optional embodiment, the motion sensing algorithm is used to calculate the motion state parameters of the target in real time, obtain the calculation results, and input them to the space-time behavior analysis unit for hierarchical response processing, including:
[0145] S631, input the calculation results to the space-time behavior analysis unit, identify the target behavior mode through the hidden Markov model, and trigger the abnormal behavior event to the early warning decision system for hierarchical response processing.
[0146] It should be noted that after the CVMS algorithm calculates the motion state parameters (speed, acceleration, and direction angle) of the target in real time, the calculation results are input to the space-time behavior analysis unit, the target behavior mode (such as wandering, running, and gathering) is identified through the hidden Markov model, and finally the abnormal behavior event is triggered to the early warning decision system for hierarchical response processing.
[0147] As shown in Figures 2-4 According to another embodiment of the present application, a target focusing device based on a six-degree-of-freedom bionic eye is also provided, which comprises:
[0148] a base 1;
[0149] a neck yaw degree of freedom 2 arranged at the top end of the base 1;
[0150] a neck connecting piece 3 arranged at the top of the neck yaw degree of freedom 2
[0151] a neck-eye connecting piece 4 movably connected to the top of the neck connecting piece 3;
[0152] an eye yaw degree of freedom 6, an eye yaw connecting piece 7, an eye pitch connecting piece 8, and an eyeball 9, which are symmetrically arranged in order from bottom to top at the top of the neck-eye connecting piece 4;
[0153] a laser radar 5 arranged at the top end of the neck-eye connecting piece 4 and located between the two groups of eyeballs 9.
[0154] It should be noted that the neck yaw degree of freedom 2, the neck connecting piece 3, the eye yaw degree of freedom 6, and the eye pitch connecting piece 8 all provide power for the autonomous adjustment of the eyeball 9 through a plurality of servo motors.
[0155] The above merely provides the preferred embodiment of the present application, and is not used to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the protection scope of the present application.
Claims
1. A panoramic perception method based on a hexapod biomimetic eye, characterized in that, The method comprises the following steps: S1, real-time acquisition of a synchronous image pair based on a binocular vision sensor, and calculation of target position information by using a binocular perception algorithm; S2, real-time scanning by using a laser radar to obtain target position cloud data for preprocessing, obtaining target point cloud coordinates based on a Euclidean clustering algorithm, and converting the target position information into target position information of the binocular vision sensor by using a coordinate system; S3, fusion of local geometric shape features in the target position information obtained by the binocular vision sensor and the laser radar respectively by using a logic block development kit to generate a shallow thermal map; fusion of global geometric shape features in the target position information obtained by the binocular vision sensor and the laser radar respectively by using an adaptive zooming algorithm to generate a deep thermal map; S4, fusion of local features in the shallow thermal map and the deep thermal map by using an attention mechanism and a feature weighting algorithm to generate a feature matrix, and calculation of a focal length by combining a lens imaging formula, real-time adjustment of the target focal length of the binocular vision sensor by using a controller, and verification of the target focal length accuracy by using a definition detection model; S5, dynamic adjustment of the interpupillary distance to match the target position information by using binocular stereo vision based on the target position information calculated by the binocular perception algorithm, real-time adjustment of the target interpupillary distance of the binocular vision sensor by using a controller, and enhancement of the interpupillary distance features by using an unsupervised reconstruction mechanism; S6, construction of a bionic motion model based on a neck-eye collaborative control algorithm, extraction of target features and motion states, and processing of the target features and the motion states by using a deep learning target detection model and a motion perception algorithm, respectively, to optimize the target position information.
2. The panoramic perception method based on a hexapod bionic eye according to claim 1, characterized in that, The method comprises the following steps: S11, construction of an iterative geometric coding cost volume by using a binocular perception algorithm to fuse the geometric shape features of the synchronous image pair; S12, multi-round optimization of the synchronous image pair by using a gated recurrent unit to calculate the distance of the target object to the binocular vision sensor.
3. The panoramic perception method based on a hexapod bionic eye according to claim 1, characterized in that, The method comprises the following steps: S21, calculation of target position original distance data by using a laser radar to emit a modulated laser beam and receive a reflected signal, and by using a time-of-flight principle; S22, removal of noise points and down-sampling processing of the obtained target position cloud data by using statistical outlier filtering and voxel grid filtering; S23, segmentation of target point cloud coordinates based on a Euclidean clustering algorithm, extraction of target point cloud clusters, and calculation of three-dimensional coordinates of cluster center points; S24, conversion of the three-dimensional coordinates of the laser radar into target position information of the binocular vision sensor by using a coordinate system.
4. The panoramic perception method based on a hexapod bionic eye according to claim 1, characterized in that, The method comprises the following steps: S3, fusion of local geometric shape features in the target position information obtained by the binocular vision sensor and the laser radar respectively by using a logic block development kit to generate a shallow thermal map; The global geometric shape features in the target position information obtained by the binocular vision sensor and the laser radar are fused by using an adaptive zoom algorithm to generate a deep heat map, including: S31, the target position information obtained by the laser radar and the target position information obtained by the binocular vision sensor are fused by using a deep learning multi-scale feature fusion method; S32, according to the local geometric shape features and the global geometric shape features in the target position information, the target position information is divided into a shallow heat map and a deep heat map.
5. The panoramic perception method based on a hexapod bionic eye according to claim 1, characterized in that, The local features in the shallow heat map and the deep heat map are fused by using an attention mechanism and a feature weighting algorithm to generate a feature matrix, and the focal length is calculated by combining the lens imaging formula, and the target focal length of the binocular vision sensor is adjusted in real time by the controller, and the target focal length accuracy is verified by a sharpness detection model, including: S41, the local features in the shallow heat map and the deep heat map are learned autonomously by using an attention mechanism, and the extracted local features are fused by using a feature weighting algorithm to generate a matrix containing target position depth and spatial position information; S42, according to the matrix after weighting fusion, the focal length is calculated in real time by combining the lens imaging formula, and the real-time calculated focal length is converted into a control signal of the controller; S43, the target focal length position of the binocular vision sensor is adjusted in real time according to the control signal of the controller; S44, the sharpness of the target position information is detected in real time by using a sharpness detection model, and the accuracy of the target focal length adjustment is further optimized.
6. The panoramic perception method based on a hexapod bionic eye according to claim 1, characterized in that, The target position information calculated based on the binocular perception algorithm is dynamically adjusted to match the target position information by using binocular stereo vision, the target pupil distance of the binocular vision sensor is adjusted in real time by the controller, and the pupil distance features are enhanced by using an unsupervised reconstruction mechanism, including: S51, generate an initial disparity map based on the binocular perception algorithm, extract target position information, dynamically adjust the pupil distance to match the target position information by combining the geometric constraints of binocular stereo vision, and adjust the target pupil distance of the binocular vision sensor in real time by the controller; S52, use the unsupervised reconstruction mechanism to construct a multi-scale feature pyramid, and enhance the weak texture region features by cyclic domain adaptation.
7. The panoramic perception method based on a hexapod bionic eye according to claim 1, characterized in that, Based on the neck-eye cooperative control algorithm, a bionic motion model is constructed, target features and motion states are extracted, and a deep learning target detection model and a motion perception algorithm are used to process the target features and motion states respectively to optimize the target position information, including: S61, based on the neck-eye cooperative control algorithm, a bionic motion model containing binocular vision information and neck motion information is constructed, and target features and motion states are extracted from the binocular vision information and neck motion information; S62, use a deep learning target detection model to detect moving target features to obtain detection box coordinates and feature vectors, and input them into a single target tracking module for processing to generate a continuous target feature sequence; S63, use a motion perception algorithm to calculate the motion state parameters of the target in real time to obtain the calculation results, and input them into a space-time behavior analysis unit for hierarchical response processing.
8. The panoramic perception method based on a hexapod bionic eye according to claim 7, characterized in that, The detection of the moving target feature by using the deep learning target detection model obtains the detection frame coordinates and the feature vector, and is input to the single target tracking module for processing to generate a continuous target feature sequence, which includes: S621, input the detection frame coordinates and the feature vector to the single target tracking module, predict the target trajectory by using the Kalman filter, and use the Hungarian algorithm for trajectory association to generate a continuous target feature sequence.
9. The panoramic perception method based on a hexapod bionic eye according to claim 8, characterized in that, The motion sensing algorithm is used to calculate the motion state parameters of the target in real time, obtain the calculation results, and input them to the space-time behavior analysis unit for hierarchical response processing, which includes: S631, input the calculation results to the space-time behavior analysis unit, identify the target behavior mode by using the hidden Markov model, and trigger the abnormal behavior event to the early warning decision system for hierarchical response processing.
10. A target focusing device based on a hexapod biomimetic eye for implementing a panoramic perception method based on a hexapod biomimetic eye according to any one of claims 1-9, characterized in that, It includes: a base; a neck yaw degree of freedom arranged at the top end of the base; a neck connecting piece arranged at the top of the neck yaw degree of freedom a neck eye connecting piece movably connected to the top of the neck connecting piece; a eyeball yaw degree of freedom, an eye yaw connecting piece, an eye pitch connecting piece and an eyeball, which are symmetrically arranged from bottom to top on the top of the neck eye connecting piece; a laser radar arranged at the top end of the neck eye connecting piece and located between the two groups of eyeballs.