Bionic robot vision coordination method based on multi-vision module and robot vision device

By employing a multi-vision module collaborative approach, the biomimetic robot system achieves collaborative perception of both global and local conditions in complex environments. This addresses the shortcomings of existing vision systems, which cannot simultaneously consider both global and local conditions, and improves the robustness and task adaptability of the vision system.

CN120839803BActive Publication Date: 2025-11-21SHANGHAI TODAY XINDONG TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511343393.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2025-11-21
Estimated Expiration
2045-09-19

AI Technical Summary

Technical Problem

Existing biomimetic robot vision systems struggle to balance morphological biomimicry with functional realization when mimicking human vision. In particular, they cannot effectively combine global and local perception in complex environments. Furthermore, multi-camera systems lack a unified coordination mechanism, leading to wasted computing resources and target features being easily overwhelmed by noise.

Method used

A multi-vision module collaborative approach is adopted. The wide-angle main view module acquires the main view image data and extracts low-level visual features. Combined with a pre-set bionic attention weight mapping table, the high-level visual feature requirements are determined, a dynamic attention heatmap is generated, and the attention area is dynamically adjusted to achieve high-resolution imaging and feature fusion, forming a perception-action closed loop.

Benefits of technology

It improves the robustness and adaptability of the robot vision system, enabling it to efficiently focus on key areas in complex environments, simulate the dynamic adjustment capabilities of human vision, and enhance the task adaptability and naturalness of human-computer interaction of bionic robots in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120839803B_ABST
    Figure CN120839803B_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to the technical field of bionic robots, and particularly relate to a bionic robot vision coordination method based on multiple vision modules, a robot vision device, a vision unit device and a robot. The method comprises: S1: receiving and analyzing task instructions to obtain task semantics; S2: obtaining main view image data and extracting low-level vision features; S3: searching for a bionic attention weight mapping table to obtain high-level vision feature requirements and initial feature weights; S4: generating an initial dynamic attention heat map to determine a high attention area and a secondary attention area; S5: obtaining high-resolution image data to update the low-level vision features of the main view image data; and S6: weighted fusion to update the dynamic attention heat map in real time to form a perception-action closed loop. The method simulates the human visual attention mechanism of top-down task-driven selective focusing and bottom-up saliency perception according to tasks, and improves the bionic degree of robot vision coordination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this application relate to the field of biomimetic robot technology, specifically to a biomimetic robot visual collaboration method based on multiple visual modules, a robot vision device, a vision unit device, and a robot. Background Technology

[0002] The statements herein are provided only as background information in connection with this application and do not necessarily constitute prior art.

[0003] With the increasing demand for bionic robots in service and companionship scenarios, enhancing the naturalness and realism of human-computer interaction has become a core requirement for technological development. By mimicking human form and behavior patterns, bionic robots aim to enhance their interactive affinity and collaborative efficiency. In this process, the visual system, as the core of environmental perception, must be designed to balance the dual goals of form biomimicry and functional realization.

[0004] Currently, there are two main technical solutions for the design of vision systems in bionic robots: one is to simulate the appearance of human eyes, focusing on morphological biomimicry. Existing solutions typically involve placing a camera in each eye. While this configuration achieves a high degree of biomimicry in appearance, its visual function differs from that of humans. The dynamic attention mechanisms of human vision (such as the coordination between foveal fine vision and peripheral vision) and the selective perception ability based on task semantics are functional characteristics that this approach has not yet fully realized. This technical solution can only achieve appearance simulation, lacking simulation in terms of function and perception modes. It cannot simultaneously achieve global and local perception, task adaptability, and natural interaction in complex environments. Secondly, improving perception performance and focusing on functional implementation are crucial. Such functional robots typically employ a combination of heterogeneous cameras, including wide-angle and telephoto lenses, to achieve complex environmental perception and task execution through engineered attention modules. However, the physical arrangement of multiple cameras cannot fully conform to the morphological constraints of a bionic head, compromising the bionic appearance. Furthermore, existing perception strategies largely rely on data-driven statistical weighting rules, lacking the ability to dynamically allocate perception resources based on task semantics. This leads to redundant consumption of computational resources in non-critical areas, and key features are easily obscured by environmental noise. Therefore, bionic robots urgently need a bionic vision method that mimics the collaborative multi-visual-module approach of human visual and attention mechanisms. Summary of the Invention

[0005] A brief overview of this application is provided below to offer a basic understanding of certain aspects thereof. It should be understood that this overview is not an exhaustive summary of the application. It is not intended to identify key or essential parts of the application, nor is it intended to limit its scope. Its purpose is merely to present certain concepts in a simplified form as a prelude to the more detailed description that follows.

[0006] In a first aspect, embodiments of this application provide a bionic robot visual collaboration method based on multiple vision modules, applied to a bionic robot equipped with multiple imaging modules. The method includes at least the following steps:

[0007] S1: Receive and parse task instructions to obtain task semantics;

[0008] S2: The wide-angle main view module synchronously acquires the main view image data and extracts low-level visual features in parallel.

[0009] S3: Based on the task semantics, find a pre-set biomimetic attention weight mapping table to obtain the high-level visual feature requirements and corresponding initial feature weights that match the task semantics.

[0010] S4: Based on the requirements of low-level visual features and high-level visual features, high-level visual features are obtained, and high-level visual features are weighted and fused according to the initial feature weights to generate an initial dynamic attention heatmap, thereby initially determining the high attention area and the secondary attention area.

[0011] S5: The functional imaging module aligns with the high-interest area and acquires its high-resolution image data, extracting its high-level visual features; at the same time, the wide-angle main view module continues to run, updating the low-level visual features of the main view image data, extracting the high-level visual features of the secondary interest area and assigning weights.

[0012] S6: Weighted fusion of high-level visual features of high-attention areas, high-level visual features of secondary-attention areas, and low-level visual features, real-time updates of dynamic attention heatmap, and dynamic adjustment of high-attention areas and secondary-attention areas to form a biomimetic perception-action closed loop.

[0013] The method provided in the embodiments of this application transforms task instructions into task semantics, converting abstract tasks into quantifiable visual requirements, thus providing task guidance for subsequent visual resource allocation. It synchronously acquires frontal view image data and extracts low-level visual features through a wide-angle main view module to mimic the global field of vision of binocular vision, achieving comprehensive environmental coverage. By searching a pre-set biomimetic attention weight mapping table, it determines the high-level visual features and initial weights required for the task semantics, mimicking the human visual attention mechanism of "adjusting focus based on the target," making the allocation of high-level visual features closer to the cognitive laws of biological vision. Based on the high-level visual feature requirements and low-level visual features, it acquires high-level visual features and weights them to generate a dynamic attention heatmap, initially dividing high-attention areas and secondary attention areas, achieving focused attention on key areas, prioritizing the processing of potential target areas in complex environments, and improving processing efficiency. It controls the functional imaging module to focus on high-level visual features. High-resolution imaging is performed on the high-attention area to extract high-level visual features, while the wide-angle main view module continuously updates the low-level visual features of the main view image data and processes the high-level visual features of the secondary attention areas. Through the collaborative mode of "global monitoring + local fine-tuning", it not only imitates the division of labor between human peripheral vision and fovea, but also takes into account the dynamic coverage of the environment and the high-precision recognition of target details, overcoming the inherent defects of a single camera that cannot take into account both the global and local aspects. By fusing the high-level visual features of the high-attention area, the high-level visual features of the secondary attention area, and the low-level visual features, the dynamic attention heatmap is updated in real time and the high-attention area and the secondary attention area are adjusted to form a perception-action closed loop. It can dynamically adjust attention according to newly acquired features, simulating the dynamic adjustment capability of human vision, so that the vision system has adaptive capability and can still efficiently focus on key areas in dynamic scenarios such as target movement and occlusion changes, thus improving the robustness of the robot vision system. This method simulates the human visual attention mechanism and solves the visual function limitations of bionic robots under morphological constraints through multi-module collaboration. It enables bionic robots to simulate the human visual attention mechanism of top-down task-driven selective focusing and bottom-up saliency perception based on the task, thereby improving the bionic degree of robot visual collaboration and enhancing the task adaptability and naturalness of human-computer interaction in complex scenarios.

[0014] Embodiments of this application also provide a robot vision device, which includes a first vision unit, a second vision unit, and two bionic eyeball shells, wherein: the two bionic eyeball shells are respectively disposed in front of the first vision unit and the second vision unit, and the surface of the bionic eyeball shell includes a first region and a second region to simulate the shape of an eyeball; the first vision unit integrates at least two imaging units, wherein the acquisition angles of at least two imaging units maintain a constant relative position and have an overlapping region after installation and fixation; the second vision unit integrates at least one imaging unit, and the acquisition angle of at least one imaging unit maintains a constant relative position and has an overlapping region with one of the imaging units of the first vision unit after installation and fixation.

[0015] Embodiments of this application also provide a vision unit device for robots, the device including an imaging module and a bionic eyeball shell, wherein: the surface of the bionic eyeball shell includes a first region and a second region to simulate the shape of an eyeball; the imaging module integrates at least two imaging units, wherein the acquisition angles of at least two imaging units maintain a constant relative position and have an overlapping area after installation and fixation.

[0016] Embodiments of this application also provide a robot, which includes any of the biomimetic robot vision devices described in this application or any vision unit device applied to a robot.

[0017] These and other advantages of this application will become more apparent from the following detailed description of preferred embodiments in conjunction with the accompanying drawings. Attached Figure Description

[0018] To further illustrate the above and other advantages and features of this application, the specific embodiments of this application will be described in more detail below with reference to the accompanying drawings. The drawings, together with the following detailed description, are included in and form a part of this specification. Elements having the same function and structure are indicated by the same reference numerals. It should be understood that these drawings only depict typical examples of this application and should not be considered as limiting the scope of this application.

[0019] Figure 1 This is a flowchart illustrating a biomimetic robot visual collaboration method based on multiple vision modules according to an embodiment of this application.

[0020] Figure 2 This is a schematic diagram of the front structure of a robot vision device according to an embodiment of this application;

[0021] Figure 3 This is a schematic diagram of a vision unit device applied to a robot according to an embodiment of this application.

[0022] It should be noted that the accompanying drawings are not necessarily drawn to scale, but are shown only in a schematic manner without affecting the reader's understanding.

[0023] Explanation of reference numerals in the attached figures:

[0024] 10. First visual unit; 20. Second visual unit; 30. Imaging unit; 40. Bionic eyeball shell; 41. First region; 42. Second region; 50. Communication interface; 60. Imaging module. Detailed Implementation

[0025] Exemplary embodiments of this application will be described below with reference to the accompanying drawings. For clarity and brevity, not all features of actual implementations are described in the specification. However, it should be understood that many implementation-specific decisions must be made in the development of any such actual embodiment to achieve the developer's specific goals, such as complying with constraints related to the system and business, and these constraints may vary depending on the implementation. Furthermore, it should be understood that while development work can be very complex and time-consuming, such development work is merely a routine task for those skilled in the art who benefit from the content of this application.

[0026] It should also be noted that, in order to avoid obscuring this application with unnecessary details, only the equipment structure and / or processing steps closely related to the solution according to this application are shown in the accompanying drawings, while other details that are not closely related to this application are omitted.

[0027] It should be noted that, unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning as understood by a person with ordinary skills in the field to which this application pertains.

[0028] In the description of the embodiments of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0029] In related technologies, the development of bionic robot vision solutions revolves around morphological simulation and functional expansion. Traditional bionic robot vision systems are mostly based on the structure of human eyes, using two isomorphic cameras to form a stereo binocular system. These systems simulate human visual organs in form by matching the optical center distance to human interpupillary distance, and mainly rely on binocular parallax calculation to obtain environmental depth information, providing the robot with basic three-dimensional spatial perception capabilities. Their camera parameters (such as field of view and resolution) are usually fixed, and in practical applications, they are mainly used for structural perception in static scenes, with limited ability to handle both large-scale environmental scanning and long-distance detail observation. With the increasing demand for complex scenarios, robots have adopted solutions integrating multiple types of cameras, such as combining imaging modules with different parameters like wide-angle, telephoto, and macro. These solutions cover different observation needs through multiple modules, but the data processing of each camera is relatively independent, lacking a unified collaborative mechanism. Furthermore, the coupling of multiple modules greatly disrupts the bionic appearance. At the same time, the system cannot dynamically adjust the processing priority and resource allocation ratio of each module according to the current task type (such as target search, dynamic following, and environmental patrol), and the collaboration and intelligence between multiple modules need to be improved.

[0030] To address the aforementioned technical problems, embodiments of this application provide a bionic robot visual collaboration method based on multiple vision modules, applicable to bionic robots equipped with multiple imaging modules. Figure 1 This is a flowchart illustrating a biomimetic robot visual collaboration method based on multiple vision modules according to an embodiment of this application, as shown below. Figure 1 As shown, the method includes at least the following steps: S1: Receive and parse task instructions to obtain task semantics; S2: The wide-angle main view module synchronously acquires main view image data and extracts low-level visual features in parallel; S3: Based on the task semantics, search a preset bionic attention weight mapping table to obtain the high-level visual feature requirements and corresponding initial feature weights that match the task semantics; S4: Based on the low-level visual features and high-level visual feature requirements, acquire high-level visual features, and perform weighted fusion of the high-level visual features according to the initial feature weights to generate an initial dynamic attention heatmap, thereby initially determining the high-attention area and the secondary attention area; S5: The functional imaging module aligns with the high-attention area and acquires its high-resolution image data, extracting its high-level visual features; simultaneously, the wide-angle main view module continues to run, updating the low-level visual features of the main view image data, extracting the high-level visual features of the secondary attention area and assigning weights; S6: Weighted fusion of the high-level visual features of the high-attention area, the high-level visual features of the secondary attention area, and the low-level visual features, updating the dynamic attention heatmap in real time, and dynamically adjusting the high-attention area and the secondary attention area to form a bionic perception-action closed loop.

[0031] The method provided in the embodiments of this application transforms task instructions into task semantics, converting abstract tasks into quantifiable visual requirements, thus providing task guidance for subsequent visual resource allocation. It synchronously acquires frontal view image data and extracts low-level visual features through a wide-angle main view module to mimic the global field of vision of binocular vision, achieving comprehensive environmental coverage. By searching a pre-set biomimetic attention weight mapping table, it determines the high-level visual features and initial weights required for the task semantics, mimicking the human visual attention mechanism of "adjusting focus based on the target," making the allocation of high-level visual features closer to the cognitive laws of biological vision. Based on the high-level visual feature requirements and low-level visual features, it acquires high-level visual features and weights them to generate a dynamic attention heatmap, initially dividing high-attention areas and secondary attention areas, achieving focused attention on key areas, prioritizing the processing of potential target areas in complex environments, and improving processing efficiency. It controls the functional imaging module to focus on high-level visual features. High-resolution imaging is performed on the high-attention area to extract high-level visual features, while the wide-angle main view module continuously updates the low-level visual features of the main view image data and processes the high-level visual features of the secondary attention areas. Through the collaborative mode of "global monitoring + local fine-tuning", it not only imitates the division of labor between human peripheral vision and fovea, but also takes into account the dynamic coverage of the environment and the high-precision recognition of target details, overcoming the inherent defects of a single camera that cannot take into account both the global and local aspects. By fusing the high-level visual features of the high-attention area, the high-level visual features of the secondary attention area, and the low-level visual features, the dynamic attention heatmap is updated in real time and the high-attention area and the secondary attention area are adjusted to form a perception-action closed loop. It can dynamically adjust attention according to newly acquired features, simulating the dynamic adjustment capability of human vision, so that the vision system has adaptive capability and can still efficiently focus on key areas in dynamic scenarios such as target movement and occlusion changes, thus improving the robustness of the robot vision system. This method simulates the human visual attention mechanism and solves the visual function limitations of bionic robots under morphological constraints through multi-module collaboration. It enables bionic robots to simulate the human visual attention mechanism of top-down task-driven selective focusing and bottom-up saliency perception based on the task, thereby improving the bionic degree of robot visual collaboration and enhancing the task adaptability and naturalness of human-computer interaction in complex scenarios.

[0032] In some embodiments, the wide-angle main view module includes two imaging modules with wide-angle imaging function. The two imaging modules are symmetrically embedded in the left and right eye sockets of the bionic robot head, and the optical center distance is set to 42~80mm (matching the human interpupillary distance range) to simulate the physiological layout of human eyes and realize the bionics of the human binocular vision system in terms of hardware form.

[0033] In some embodiments, a functional imaging module refers to an imaging component that can provide image data with higher resolution than a wide-angle main view module or specific imaging functions (such as telephoto imaging, macro imaging, infrared imaging, etc.) for areas of high interest, and whose orientation can be adjusted according to task requirements.

[0034] In some embodiments, task semantics includes at least the target task, the target task sequence, the task priority, and the target object.

[0035] The embodiments provided in this application parse the target task, target task sequence, task priority, and target object from the task instructions, enabling the bionic robot to clearly understand the task instructions, task execution order, resource allocation priority, and specific objects of visual attention of the vision module. This allows the imaging module to be scheduled to collaboratively complete parallel tasks according to the task content, effectively improving the adaptability of the bionic robot's visual collaboration in complex task scenarios.

[0036] In some embodiments, task instructions can be parsed using a pre-trained language model, such as the BERT (Bidirectional Encoder Representations from Transformers) model.

[0037] In some embodiments, low-level visual features include edges, corners, and depth.

[0038] The embodiments provided in this application extract edges, corners, and depth from front-view image data to simulate overall perception in human vision. This provides structured data for generating dynamic attention heatmaps, ensuring that the division of high-attention areas more closely matches the structure of real scenes, reducing deviations caused by missing underlying features, and improving the robustness of basic perception in visual collaboration. Specifically, edge features can quickly identify object outlines and scene boundaries, providing a basis for locating potential occlusions and corner areas; corner features can mark turning points in spatial structures, assisting in constructing the spatial layout of the scene; and depth features generate three-dimensional distance information through binocular parallax, accurately distinguishing the front-back positional relationships of objects.

[0039] In some embodiments, advanced visual features are features related to the target object, including the target object's geometry, spatial location, texture, semantic labels, optical character recognition (OCR) regions, and motion state.

[0040] The embodiments provided in this application acquire advanced visual features related to the target object, such as geometric shape, spatial location, texture, semantic tags, optical character recognition (OCR) regions, and motion state, which can accurately depict the multi-dimensional visual attributes of the target object. Among them, geometric shape features can distinguish the target object from interference objects; spatial location features combined with depth information determine the three-dimensional coordinates of the target, which can guide the steering control of the functional imaging module; texture features can enhance the identification of details of the target object; semantic tags are directly related to task semantics; OCR region features can parse the text information attached to the target object; and motion state features can support dynamic target tracking. The synergistic effect of these advanced features not only meets the differentiated needs of different tasks for target description, but also provides a basis for the weight adjustment of dynamic attention heatmaps, improving the accuracy of target recognition in complex scenes, the stability of dynamic tracking, and the flexibility of task adaptation, and is closer to the human visual perception of targets.

[0041] In some embodiments, the dynamic attention heatmap and the biomimetic attention weight mapping table are both constructed based on the Itti-Koch model.

[0042] The embodiments provided in this application construct a dynamic attention heatmap and a biomimetic attention weight mapping table based on the Itti-Koch model. Based on the biomimetic core of biological vision using the Itti-Koch model, and through the combination with task semantic parsing and multi-module collaboration, the attention regulation mechanism of the visual system is made closer to the cognitive laws of human vision. This achieves dual-pathway collaboration of bottom-up saliency perception and top-down task-driven approaches, improving the robustness of task execution and the naturalness of human-computer interaction in complex scenarios.

[0043] In some embodiments, when constructing a biomimetic attention weight mapping table based on the Itti-Koch model, a feature set can be formed by combining the original feature channels of the Itti-Koch model with the extracted visual feature types. ,in,

[0044] Low-level visual features This represents advanced visual features. Depth features, as an extended dimension, are incorporated into saliency calculations through their correlation with brightness and edge features.

[0045] Using the parsed task semantics as the core elements, an association matrix is ​​constructed. , where, element This represents the association strength between the m-th task semantic and the i-th visual feature, for example:

[0046] If the feature is a core identifying feature of the target object (such as "green plants" and "green texture"), then =0.8~1.0;

[0047] If the features are key supporting features of the target object (such as "search task" and "corner spatial location"), then =0.6~0.9;

[0048] If the feature is strongly correlated with task priority and the time sequence of the target task (e.g., "high-priority task" and "target motion state"), then or =0.5~0.8;

[0049] If there is no direct connection, then =0.

[0050] By combining the bottom-up saliency of the Itti-Koch model with the top-down modulation of task semantics, the feature weights conform to formula (1):

[0051] (1).

[0052] in, Features output by the Itti-Koch model Baseline saliency value (normalized to [0,1], low-level visual features primarily depend on this value); {T}_{m}\in [0,1] For the first m The priority coefficient of each task semantics (determined by task semantic parsing, such as high-priority tasks). );\alpha \in [0,1] The balance coefficient (the initial scanning phase focuses on low-level visual features), The focusing phase emphasizes high-level visual features. ).

[0053] The correlation matrix is ​​dynamically adjusted based on the task execution phase (global scan - local focus - target confirmation). With balance coefficient This forms a stage sub-table, where each stage adapts to different stages of task execution (such as global scanning, local focusing, target confirmation, etc.). The biomimetic attention weight mapping table contains a pre-defined set of detailed rules for assigning feature weights to each stage. For example:

[0054] Global scanning: Enhances the weight of low-level visual features such as edges and corners, as well as high-level visual features such as spatial location. ;

[0055] Local focusing phase: Increase the weight of high-level visual features such as the geometry and texture of the target object. ;

[0056] Target confirmation phase: Maximize the weights of core high-level visual features such as semantic labels and OCR regions. .

[0057] In some embodiments, when generating a dynamic attention heatmap based on the Itti-Koch model, the main view image can be divided into grid cells of a preset size, and the feature response value within each cell can be calculated. For example, the low-level visual feature response value can be the normalized result of edge density and depth gradient within the cell (value 0-1); the high-level visual feature response value can be the normalized result of target shape matching degree and texture similarity within the cell (value 0-1).

[0058] For each grid cell, based on the weights obtained from the pre-set bionic attention weight mapping table, the low-level visual feature response values ​​and the high-level visual feature response values ​​are weighted and summed to obtain the cell attention level. The mathematical relationship between them conforms to formula (2):

[0059] (2).

[0060] in, For grid cells Unit attention, For low-level visual feature weights, These are the low-level characteristic response values ​​within the unit. For high-level visual feature weights, This refers to the high-level characteristic response value within the cell.

[0061] By setting a first attention threshold and a second attention threshold in conjunction with the task phase, continuous grid cells with an attention level not lower than the first attention threshold are classified as high attention regions, and continuous grid cells with an attention level lower than the first attention threshold but not lower than the second attention threshold are classified as secondary attention regions. Gaussian filtering is used to spatially smooth the distribution of unit attention levels, eliminating regional fragmentation caused by isolated high-response cells, and finally outputting a dynamic attention heatmap with continuous high attention regions.

[0062] In some embodiments, step S5 further includes the following steps: S51: Verify whether the high-level visual features of the high-attention region meet the requirements of the task semantics; if not, the region is determined to be a misjudged region; S52: Reduce the weight of the misjudged region in the dynamic attention heatmap; S53: Select the region with the highest feature weight from the current secondary attention regions as the new high-attention region; S54: Control the functional imaging module to turn to the new high-attention region and acquire high-resolution image data; S55: Repeat steps S51 to S54 until the region containing the high-level visual features required by the task semantics is successfully located or all secondary attention regions have been traversed; S56: If the region that meets the requirements of the task semantics is still not successfully located after traversing all secondary attention regions, the dynamic attention heatmap is regenerated based on the latest main view image data and task semantics.

[0063] The embodiments provided in this application achieve the effects of avoiding invalid visual computation and significantly shortening the target localization time by verifying whether the high-level visual features of the high-interest region match the task semantics, reducing the weight of misjudged regions, and selecting the region with the highest feature weight from the secondary interest region for relocalization (instead of returning to global scanning). By shifting the functional imaging module to the new high-interest region after discovering the misjudged region while retaining the wide-angle main view module to continuously update image data and features, visual perception discontinuity is avoided and the collaborative continuity of multiple modules is ensured. Furthermore, by regenerating a dynamic attention heatmap based on the latest main view image and task semantics when the secondary region is still not located, a mechanism of local correction and global update collaboration is formed, which avoids task interruption due to misjudgment and improves the robustness of task execution in complex environments. In the actual scenario of bionic robot applications, this method can enable robots to reduce invalid visual exploration, avoid being stuck due to misjudged targets, and always advance the task coherently and efficiently, making it more suitable for the complex needs of real environments with many interferences and possible target movement.

[0064] In some embodiments, step S6 further includes the following steps: if a new task instruction with a higher task priority is received, the dynamic attention heatmap is updated based on the task semantics of the new task instruction, and the high attention region of the original task instruction is set as the secondary attention region of the new task instruction.

[0065] The embodiments provided in this application achieve rapid adaptation to task switching by prioritizing responses to new tasks with higher priority and updating the dynamic attention heatmap in real time based on the semantics of the new task. This avoids response delays caused by task priority conflicts, improves the robot's efficiency in responding to dynamic task demands, and achieves efficient reuse of visual resources by setting the high-attention area of ​​the original task as the secondary attention area of ​​the new task, instead of directly discarding the visual data of the original task. After the new task is completed, the attention points of the original task can be quickly traced back without re-scanning the global system, reducing unnecessary calculations. In practical application scenarios such as home services (e.g., switching from "tidying up the desktop" to "handling spilled food"), this design allows the bionic robot to prioritize handling urgent / high-value tasks while retaining the progress of the original task, meeting the needs of dynamic alternation of multiple tasks in real-world scenarios, improving the robot's practical adaptability, and making the robot's behavior closer to that of humans.

[0066] Embodiments of this application also provide a robot vision device. Figure 2 This is a schematic diagram of the front structure of a robot vision device according to an embodiment of this application, as shown below. Figure 2 As shown, the device includes a first visual unit 10, a second visual unit 20, and two bionic eyeball shells 40, wherein: the two bionic eyeball shells 40 are respectively disposed in front of the first visual unit 10 and the second visual unit 20, and the surface of the bionic eyeball shell 40 includes a first region 41 and a second region 42 to simulate the shape of an eyeball; the first visual unit 10 integrates at least two imaging units 30, wherein the acquisition angles of at least two imaging units 30 maintain a constant relative position and have an overlapping area after installation and fixation; the second visual unit 20 integrates at least one imaging unit 30, and the acquisition angle of at least one imaging unit 30 maintains a constant relative position and has an overlapping area with one of the imaging units 30 of the first visual unit 10 after installation and fixation.

[0067] The device provided in the embodiments of this application, by placing the bionic eyeball shell 40 in front of the first visual unit 10 and the second visual unit 20, allows the robot vision device to closely resemble the shape of an eye in appearance, improving visual affinity during human-computer interaction and reducing the sense of alienation when humans interact with robots. At the same time, by setting multiple imaging units 30 in the first visual unit 10 and the second visual unit 20, a structured layout of at least three imaging units 30 is achieved, providing hardware support for robot visual perception. The constant overlapping viewpoint of at least two imaging units 30 in the first visual unit 10 can stably capture multi-dimensional image information of the same scene and avoid feature misalignment caused by viewpoint shift. The overlapping viewpoint of the second visual unit 20 with the imaging unit 30 of the first visual unit 10 further expands the effective perception range while ensuring the spatial correlation of multi-source image data, which facilitates subsequent image data processing. This device addresses the issues of limited binocular vision in existing bionic robots and the lack of bionic design in functional robot vision systems. It improves the overall performance of bionic robot vision devices, enabling them to facilitate close-range human-robot interaction by mimicking human eyes in design, thus reducing the mechanical feel of human-robot interactions. Furthermore, it leverages the structured layout of multiple imaging units and a constant overlapping field of view to stably output multi-dimensional, spatially correlated image data, providing reliable hardware support for subsequent feature extraction and target recognition. This allows bionic robots to achieve accurate visual perception even in complex environments, enhancing their adaptability and practical value in real-world applications.

[0068] In some embodiments, the bionic eyeball shell 40 can be any structure that can present the appearance of an eye, such as a closed spherical shell, an open spherical shell, a hemispherical shell, a streamlined shell, and a thick spindle-shaped object, etc., and this application does not limit it.

[0069] In some embodiments, the bionic eyeball shell 40 may be a closed structure, with the first visual unit 10 and the second visual unit 20 disposed inside their respective bionic eyeball shells 40 and communicating with an external processor.

[0070] In some embodiments, this application does not limit the center-to-center distance of the first region 41 of the two bionic eyeball shells 40, but preferably, the center-to-center distance of the first region 41 of the two bionic eyeball shells 40 can be set to 42~80mm to mimic the common range of human eye distance.

[0071] In some embodiments, the first region 41 is a circular region that simulates the appearance of the pupil iris (black of the eye), and the second region 42 is a region surrounding the first region 41 that simulates the appearance of the sclera (white of the eye). Both the first region 41 and the second region 42 are translucent along the light transmission path of the imaging unit 30. The embodiments of this application can both restore the shape of the human eye through the partitioning of the bionic eyeball shell 40 and ensure that light can smoothly enter the imaging unit 30, thus achieving compatibility between the bionic appearance and the visual acquisition function.

[0072] In some embodiments, the specific shape and material of the first region 41 can be flexibly set: for example, it can be a through hole, an opaque material, a transparent circular material, or a dark-colored light-transmitting material with a light transmittance of more than 85%, such as stained optical glass, dark optical acrylic, dark structural optical film, etc. This application does not limit this.

[0073] In some embodiments, the material of the second region 42 can also be flexibly selected, for example, it can be a light-colored opaque material or a light-transmitting light-colored material, such as milky white light-transmitting PC (polycarbonate), light-colored optical silicone, etc. This application does not limit this.

[0074] In some embodiments, the first visual unit 10 and the second visual unit 20 may further include a feature extraction module configured to perform low-level feature extraction on the image acquired by the imaging unit 30.

[0075] In some embodiments, the device further includes a communication interface 50 for transmitting image data acquired by the imaging unit 30 to a processor for processing.

[0076] The embodiments provided in this application efficiently and stably transmit image data to the processor through the communication interface 50, avoiding data loss or delay. This ensures the accuracy of subsequent image processing and provides timely data support for the adaptive adjustment of the imaging unit 30, which is beneficial for the robot to accurately complete tasks such as tracking and detection.

[0077] In some embodiments, the first region 41 is configured to simulate the appearance of the iris region of the human eye pupil, and the second region 42 is configured to simulate the appearance of the sclera region of the human eye.

[0078] The embodiments provided in this application, by setting the first region 41 and the second region 42 to resemble the appearance of the iris region of an adult eye and the sclera region of a human eye respectively, can highly reproduce the appearance structure of the human eye, making the robot vision device visually closer to the shape of a real human eye, significantly improving the appearance simulation, which helps to weaken the mechanical feel of the robot and reduce the psychological alienation when humans come into contact with the robot. At the same time, the region division that conforms to the human eye's cognitive habits also makes the robot's "visual expression" easier for humans to understand, further strengthening the practical value of bionic design.

[0079] In some embodiments, the first visual unit 10 and the second visual unit 20 each include at least one wide-angle unit that implements wide-angle imaging function, and among the plurality of imaging units 30, in addition to the two wide-angle units, at least one functional imaging unit is also included.

[0080] The embodiments provided in this application simulate the peripheral vision of human eyes through two wide-angle units, supporting binocular imaging and global environmental perception. Combined with a functional imaging unit to mimic the fine recognition capabilities of central vision, it helps to accurately extract features during the local focusing stage, forming a global-local collaborative closed loop, which not only improves visual biomimicry but also enhances recognition accuracy and efficiency.

[0081] In some embodiments, a functional imaging unit with specific imaging capabilities may be selected according to the specific application scenario of the bionic robot. The imaging capability may be one or more of telephoto imaging, macro imaging, motion imaging, and infrared imaging.

[0082] In some embodiments, at least two imaging units 30 in the first visual unit 10 have parallel optical axes and a spacing of less than or equal to 16 mm; when there are multiple imaging units 30 in the second visual unit 20, at least two imaging units 30 have parallel optical axes and a spacing of less than or equal to 16 mm.

[0083] The embodiments provided in this application achieve the purpose of adapting multiple imaging units 30 to the limited internal space of the bionic eyeball shell 40 by making the optical axes of at least two imaging units 30 within the first visual unit 10 or the second visual unit 20 parallel and the distance between them less than or equal to 16mm. This avoids exceeding the size limit of the human eyeball shape due to the excessively large distribution distance of multiple imaging units 30, and enables the imaging units 30 to be stably arranged within the bionic structure of the bionic eyeball. This ensures that the bionic appearance of the eyeball is not damaged when integrating multiple imaging units 30, and makes the compactness of the functional units and the bionic form compatible, which is more in line with the assembly requirements of the bionic human eyeball.

[0084] Embodiments of this application also provide a vision unit device for robots. Figure 3 This is a structural schematic diagram of a vision unit device applied to a robot according to an embodiment of this application, such as... Figure 3 As shown, the device includes an imaging module 60 and a bionic eyeball shell 40, wherein: the surface of the bionic eyeball shell 40 includes a first region 41 and a second region 42 to simulate the shape of an eyeball; the imaging module 60 integrates at least two imaging units 30, wherein the acquisition angles of at least two imaging units 30 maintain a constant relative position and have an overlapping area after installation and fixation.

[0085] In the embodiments provided in this application, by placing the bionic eyeball shell 40 in front of the imaging module 60, the robot vision device can closely resemble the shape of an eye in appearance, thereby improving visual affinity during human-computer interaction and reducing the sense of alienation when humans interact with robots. At the same time, by setting multiple imaging units 30, multimodal hardware support is provided for robot visual perception. Among them, the constant overlapping viewpoint of at least two imaging units 30 can stably capture multi-dimensional image information of the same scene and avoid feature misalignment caused by viewpoint shift. This device can achieve the purpose of bionic appearance to adapt to human-computer interaction scenarios and multi-vision modules working together to ensure visual perception.

[0086] In some embodiments, the plurality of imaging units 30 are configured to perform at least two imaging functions.

[0087] The embodiments provided in this application, by configuring multiple imaging units 30 to achieve at least two imaging functions, avoid the limitations of the single imaging function of traditional robot vision systems, enhance the overall adaptability and flexibility of the vision system, and help improve the robot's ability to process diverse visual perception tasks.

[0088] In some embodiments, the imaging function of the imaging unit 30 may be one or more of wide-angle imaging, telephoto imaging, macro imaging, motion imaging, and infrared imaging.

[0089] Embodiments of this application also provide a robot, which includes any of the biomimetic robot vision devices described in this application or any vision unit device applied to a robot.

[0090] The embodiments provided in this application enhance visual affinity during human-computer interaction and reduce the sense of alienation when humans interact with robots by incorporating any of the biomimetic robot vision devices or vision unit devices applied to robots in this application. At the same time, they provide multimodal hardware support for robot visual perception, thereby achieving the goal of biomimetic robot appearance to adapt to human-computer interaction scenarios and multi-vision module collaboration to ensure visual perception.

[0091] The following example, using the task instruction "Put the green plant on the table on the ground," illustrates in detail how the inventors supplement one or more embodiments mentioned above using the robot vision device and the bionic robot vision collaboration method based on multiple vision modules provided in this application.

[0092] The bionic eyeball shell consists of two hemispherical shells with a diameter of 35mm. The first area is made of dark-colored transparent glass with a diameter of 8mm (90% light transmittance), and the second area is made of milky white transparent PC material. The center point of the first area of ​​the two shells is 65mm apart, and they are respectively placed in front of the first visual unit and the second visual unit. The visual unit is embedded inside the shell.

[0093] The first visual unit integrates two imaging units. The wide-angle unit has a field of view of 120° and a resolution of 1920×1080. The functional imaging unit is a telephoto lens that supports 4K resolution. The optical axes of the two units are parallel and the distance between them is 12mm. The overlap rate of the acquisition angle is ≥70%.

[0094] The second visual unit integrates a wide-angle unit with parameters identical to those of the wide-angle unit of the first visual unit. Its acquisition viewpoint overlaps with the wide-angle unit of the first visual unit by ≥80%, ensuring comprehensive scene coverage without blind spots.

[0095] The imaging unit can be a camera such as the Xiaomi Civi series, and the communication interface adopts a high-speed data interface to transmit the image data collected by the imaging unit to the robot processor in real time.

[0096] The task instruction to "place the green plant on the table on the ground" is executed as follows:

[0097] The robot receives the user's command "Put the green plant on the table on the ground" through the voice interaction module. The processor calls the BERT model to parse the task semantics and obtain the following task semantics:

[0098] Objective: Pick up the green plant on the table and place it on the ground;

[0099] Target: Green plants on the table (Characteristics: with leaves, with a flowerpot, located in the desktop area);

[0100] Task priority: Normal priority, There are no higher priority tasks;

[0101] Task sequence: No subtasks in any order, directly execute the location-grab-place sequence.

[0102] The wide-angle units in the first and second visual units are activated simultaneously to acquire real-time global images of the living room (30 FPS). The processor extracts low-level visual features in parallel, obtaining the following low-level visual features:

[0103] Edge features: The desktop outline (rectangular edge), plant leaf edge (irregular curve), and flower pot edge (circular or square) are extracted using the Canny algorithm.

[0104] Corner feature: The four corners of the table and the corners where the flowerpot base contacts the tabletop are marked using the Harris algorithm;

[0105] Depth features: Based on binocular parallax calculation, the height of the top of the plant from the table (approximately 20cm), the height of the table from the ground (approximately 75cm), and the current distance between the plant and the robot (approximately 1.5m) are obtained.

[0106] Based on the task semantics, the processor calls a pre-built biomimetic attention weight mapping table constructed based on the Itti-Koch model to determine the following:

[0107] Advanced visual feature requirements: the geometric shape of the green plant (feathered / oval), spatial position (within the desktop area), texture (leaf vein texture), and motion state (static, no movement);

[0108] Initial feature weights: based on the correlation matrix Calculate (M=4, 4 task semantic elements; n=6, 6 high-level visual features and low-level visual features).

[0109] Green plant texture (core identifying feature): The target object is strongly correlated with the features;

[0110] Green plant geometry (key supporting features): Distinguish between green plants and desktop clutter;

[0111] Green space location (desktop area): Define the target scope;

[0112] Other characteristics (such as motion state): Unrelated;

[0113] Balance coefficient Initial scanning phase It focuses on low-level features to ensure global positioning.

[0114] The processor acquires the high-level visual features of the current scene based on low-level visual features (edges, corners, depth) and high-level visual feature requirements, and calculates the unit attention of the image grid unit using formula (2), with each unit consisting of 10×10 pixels:

[0115] For desktop area grid cells: low-level feature response values (Clear edges and corners), high-level feature response values (There is a suspected green plant texture), weight , Unit attention =0.8×0.7+0.6×0.3=0.74;

[0116] For non-desktop areas (such as walls and floors): (Blurred edges) (No green plant features) =0.2×0.7+0×0.3=0.14;

[0117] Set the first attention threshold to 0.6 and the second attention threshold to 0.3, and define the high-attention area as the area within the desktop. For continuous grid cells ≥0.6, the secondary interest area is defined as 0.3≤ on the desktop. A dynamic attention heatmap is generated using grid cells with a resolution of <0.6. A 5×5 Gaussian filter is then applied to the dynamic attention heatmap to eliminate isolated high-response cells caused by desktop glare, ensuring that the high-attention area is continuous and free of fragmentation.

[0118] The processor controls the functional imaging unit of the first visual unit to turn towards the area of ​​high interest, acquire a 4K high-resolution image of the greenery, and extract the following high-level visual features:

[0119] Geometric shape: confirm that the leaves are oval and the flowerpot is cylindrical (excluding the square paper towel box on the table).

[0120] Texture: The reticulated veins of the leaves are clearly visible (confirming it is a green plant, not a fake plastic plant).

[0121] If a box with a green sticker (its edges and color are similar to green plants) exists on the desktop, the functional unit extracts features and finds that there is no leaf vein texture, and determines it to be a misjudged area.

[0122] From the secondary focus areas, select the "left side of the desktop with the leaf outline" as the new high focus area, which has the highest weight.

[0123] The functional imaging unit refocused on the area, confirmed it was the target green plant, and completed the localization.

[0124] The wide-angle unit is always running, updating in real time whether the desktop is obstructed (such as when the user reaches for something), and maintaining the feature weights of secondary attention areas.

[0125] By integrating the following three types of features, the unit attention is recalculated to update the dynamic attention heatmap. Emphasis on advanced features:

[0126] High-attention areas: Advanced visual features (vegetable textures, shapes): ;

[0127] Secondary focus areas: High-level visual features of the desktop edge. ;

[0128] Low-level visual features (the robot moves to a distance of 0.8m from the greenery, and the depth value is updated to 0.8m).

[0129] In the updated dynamic attention heatmap, the attention score of the grid cell containing the green plants has increased to 0.92, remaining a high-attention area.

[0130] Based on the updated heatmap, the robot controls the robotic arm to aim at the green plants in the high-attention area, and adjusts the grasping force in combination with depth data (to avoid crushing the leaves); after grasping, the wide-angle unit tracks the position of the green plants in real time, and the dynamic attention heatmap is updated as the green plants move (the high-attention area moves from the table to the end of the robotic arm).

[0131] After the robotic arm places the plant on the ground, the wide-angle unit confirms that the plant is on the ground (the relative depth value between the plant and the ground is updated to 0, which conforms to the task semantics of "placed on the ground"). The processor determines that the task is completed and turns off the high-load mode of the functional imaging unit (only the wide-angle unit monitoring is retained).

[0132] Regarding the embodiments of this application, it should also be noted that, without conflict, the embodiments of this application and the features in the embodiments can be combined with each other to obtain new embodiments.

[0133] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. The scope of protection of this application shall be determined by the scope of the claims.

Claims

1. A bionic robot visual collaboration method based on multiple vision modules, applied to a bionic robot equipped with multiple imaging modules, characterized in that, The method includes at least the following steps: S1: Receive and parse task instructions to obtain task semantics; S2: The wide-angle main view module synchronously acquires the main view image data and extracts low-level visual features in parallel. S3: Based on the task semantics, find a pre-set bionic attention weight mapping table to obtain the high-level visual feature requirements and corresponding initial feature weights that match the task semantics. S4: Based on the low-level visual features and the high-level visual feature requirements, obtain high-level visual features, and perform weighted fusion of the high-level visual features according to the initial feature weights to generate an initial dynamic attention heatmap, thereby initially determining the high attention region and the secondary attention region. S5: The functional imaging module aligns with the high-interest area and acquires its high-resolution image data, extracting its high-level visual features; at the same time, the wide-angle main view module continues to run, updating the low-level visual features of the main view image data, extracting the high-level visual features of the secondary interest area and assigning weights. S6: Weighted fusion of the high-level visual features of the high-attention region, the high-level visual features of the secondary-attention region, and the low-level visual features, real-time updates of the dynamic attention heatmap, and dynamic adjustment of the high-attention region and the secondary-attention region to form a biomimetic perception-action closed loop.

2. The method as described in claim 1, characterized in that, The task semantics include at least the target task, the target task sequence, the task priority, and the target object.

3. The method as described in claim 1, characterized in that, The low-level visual features include edges, corners, and depth.

4. The method as described in claim 2, characterized in that, The advanced visual features are features related to the target object, including the target object's geometry, spatial location, texture, semantic tags, optical character recognition (OCR) region, and motion state.

5. The method as described in claim 1, characterized in that, The dynamic attention heatmap and the biomimetic attention weight mapping table are both constructed based on the Itti-Koch model.

6. The method as described in claim 1, characterized in that, Step S5 also includes the following steps: S51: Verify whether the high-level visual features of the high-attention region meet the requirements of the task semantics. If not, the region is determined to be a misjudged region. S52: Reduce the weight of the misjudged region in the dynamic attention heatmap; S53: Select the region with the highest feature weight from the currently mentioned secondary regions of interest as the new high-interest region; S54: Control the functional imaging module to turn to the new high-interest area and acquire the high-resolution image data; S55: Repeat steps S51 to S54 until the region containing the high-level visual features required by the task semantics is successfully located or all the secondary areas of interest have been traversed. S56: If, after traversing all the secondary attention regions, the region that meets the requirements of the task semantics is still not successfully located, then the dynamic attention heatmap is regenerated based on the latest main view image data and the task semantics.

7. The method as described in claim 1, characterized in that, Step S6 also includes the following steps: If a new task instruction with a higher priority is received, the dynamic attention heatmap is updated based on the task semantics of the new task instruction, and the high-attention region of the original task instruction is set as the secondary attention region of the new task instruction.

8. A robot vision device for implementing the method as described in claim 1, characterized in that, The device includes a first visual unit, a second visual unit, and two bionic eyeball shells, wherein: The two bionic eyeball shells are respectively disposed in front of the first visual unit and the second visual unit, and the surface of the bionic eyeball shell includes a first region and a second region to simulate the shape of an eyeball; The first vision unit integrates at least two imaging units, wherein the acquisition angles of at least two of the imaging units maintain a constant relative position and have an overlapping area after installation and fixing; The second vision unit integrates at least one of the imaging units, and the acquisition angle of at least one of the imaging units maintains a constant relative position and has an overlapping area with one of the imaging units of the first vision unit after installation and fixing. The device also includes a communication interface for transmitting image data acquired by the imaging unit to a processor for processing. The first visual unit and the second visual unit each include at least one wide-angle main view module that realizes wide-angle imaging function, and among the plurality of imaging units, in addition to the two wide-angle main view modules, at least one functional imaging module is also included.

9. The apparatus as described in claim 8, characterized in that, The first region is set to simulate the appearance of the iris region of the human eye pupil, and the second region is set to simulate the appearance of the sclera region of the human eye.

10. The apparatus as claimed in claim 8, characterized in that, In the first visual unit, at least two of the imaging units have parallel optical axes and a spacing of less than or equal to 16 mm; in the second visual unit, when there are multiple imaging units, at least two of the imaging units have parallel optical axes and a spacing of less than or equal to 16 mm.

11. A robot, characterized in that, Includes the apparatus as described in any one of claims 8-10.

Citation Information

Patent Citations

  • Mechanical arm hole forming and positioning system based on laser alignment and visual measurement

    CN110539309A

  • Drug image recognition method based on artificial intelligence

    CN120496042A