An AI vision-based bionic robot control method and system

CN122645274APending Publication Date: 2026-08-28JINGHUA YUSHUI TECHNOLOGY (SUZHOU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610456669.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-08
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

[0003]采集到的作业场景视觉数据易受机器人本体振动、姿态变化、运动模糊和光照波动影响,导致目标边界、障碍区域及可操作区域识别不稳定

Benefits of technology

[0088]This invention synchronously collects and spatiotemporally aligns visual data, body motion state data, and foot contact state data of a biomimetic robot's operational scenario. In the visual preprocessing stage, it introduces motion compensation based on body motion state, region enhancement for the current control task, and a key region temporal alignment mechanism, resulting in higher stability and control relevance of the standard visual control data entering the recognition stage. In the AI ​​visual recognition stage, an improved Grounding DINO network, incorporating state-constrained attention, dual-granularity evolutionary query, and gated temporal memory update, is used to identify, locate, and extract control visual features from target objects, obstacles, and task-related regions. This ensures the output is simultaneously constrained by body state, task semantics, and temporal continuity information, thereby improving target recognition accuracy, spatial relationship analysis capability, and control constraint representation capability. In the control state construction stage, the control visual feature parameter set is fused with body motion state data and foot contact state data through control state relationship analysis and constraint information fusion, forming a control state vector containing target state, environmental constraint state, and body execution state. This enhances the global consistency and execution targeting of subsequent hierarchical decision-making and control planning. During the execution control phase, the execution control parameters are parsed into control commands through inverse kinematics solving and trajectory interpolation methods. Error analysis and closed-loop control correction are then performed using execution feedback data, forming an integrated closed-loop control mechanism encompassing perception, analysis, planning, execution, and feedback. This reduces control deviations caused by visual jitter, state mismatch, and environmental disturbances, improving the bionic robot's environmental adaptability, motion planning accuracy, real-time control performance, and task execution stability in complex scenarios. Therefore, this invention is of great significance for improving the autonomous control level of bionic robots, reducing task failure rates, and enhancing reliable operation under complex working conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122645274A_ABST
    Figure CN122645274A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on AI vision's bionic robot control method and system, including step one: synchronous acquisition bionic robot in the visual perception area in the working scene visual data, ontology motion state data and foot end contact state data;Step two: working scene visual data is preprocessed to the vision of control task;Step three: by improved Grounding DINO network, AI visual recognition and control visual feature extraction are carried out to standard visual control data;Step four: control state relationship analysis and constraint information fusion are executed;Step five: hierarchical decision analysis and control planning are carried out, and execution control parameter is obtained;Step six: execution control parameter is analyzed into control instruction;Step seven: error analysis and closed-loop control correction are carried out by collecting execution feedback data.This application improves the control task execution capability of bionic robot by improved Grounding DINO network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomimetic robot intelligent control technology, and in particular to a biomimetic robot control method and system based on AI vision. Background Technology

[0002] With the increasing demand for autonomous control in biomimetic robots, embodied intelligent systems, and complex work scenarios, multi-source information fusion technology for visual perception, body motion state analysis, and autonomous closed-loop control in robot work scenarios has received widespread attention. Existing biomimetic robot control methods mainly rely on preset rules, single sensor feedback, or ordinary visual recognition models for target perception and motion control. However, these methods generally suffer from the following problems in practical applications:

[0003] The collected visual data of the operational scene is easily affected by robot vibration, posture changes, motion blur, and lighting fluctuations, leading to unstable recognition of target boundaries, obstacle areas, and operable areas. Visual data, body motion state data, and foot contact state data often come from different sources, have inconsistent sampling frequencies, and lack a unified time reference. Existing multi-source alignment and synchronization processing technologies often struggle to achieve accurate temporal correlation and spatial correspondence across modal data, resulting in lag in state perception and deviations in control response. For the extraction of control visual features in dynamic environments, traditional target detection networks or general visual models typically focus only on target category recognition and position detection, lacking the ability to model state constraints, perform temporal query evolution, and decode control semantics for robot control tasks. This makes it difficult for the extracted results to directly represent the target's relative orientation, obstacle constraint relationships, and executable area information, affecting the accuracy of control state construction and the real-time performance of decision planning. Existing biomimetic robot control schemes often lack effective state relationship parsing and constraint information fusion mechanisms between perception results and execution control, making it difficult to simultaneously consider the coupling relationship between target state, environmental constraint state, and body execution state. This easily leads to problems such as discontinuous motion planning, unstable obstacle avoidance decisions, and inaccurate posture adjustment.

[0004] Therefore, how to provide a biomimetic robot control method and system based on AI vision is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] One objective of this invention is to propose a biomimetic robot control method and system based on AI vision. This invention constructs a target recognition, state perception, motion planning, and drive control process for a biomimetic robot through visual preprocessing for control tasks, AI visual recognition and control feature extraction using an improved Grounding DINO network, control state relationship analysis, and hierarchical decision planning. This approach offers advantages such as strong environmental adaptability, high control precision, good real-time response, and high task execution stability.

[0006] A biomimetic robot control method based on AI vision according to an embodiment of the present invention includes the following steps:

[0007] Step 1: Simultaneously collect visual data of the working scene, motion state data of the body, and contact state data of the foot end of the bionic robot within the visual perception area;

[0008] Step 2: Perform visual preprocessing on the visual data of the work scene to obtain standard visual control data;

[0009] Step 3: Using an improved Grounding DINO network, AI visual recognition and control visual feature extraction are performed on standard visual control data to obtain the control visual feature parameter set of the bionic robot at the current moment; the improved Grounding DINO network introduces a state-constrained attention mechanism, dual-granularity evolution query, and gated temporal memory update;

[0010] Step 4: Perform control state relationship analysis and constraint information fusion on the control visual feature parameter set, body motion state data and foot contact state data to obtain the control state vector of the bionic robot at the current moment;

[0011] Step 5: Based on the control state vector, perform hierarchical decision analysis and control planning to obtain the execution control parameters of each joint drive unit, foot actuator and end effector of the bionic robot;

[0012] Step 6: The execution control parameters are parsed into control commands through inverse kinematics solution and trajectory interpolation, and the bionic robot is driven and controlled based on the control commands;

[0013] Step 7: During the process of the bionic robot performing control tasks, collect execution feedback data in real time, and perform error analysis and closed-loop control correction based on the execution feedback data.

[0014] Optionally, step one specifically includes:

[0015] Obtain the visual perception center direction, visual perception range, and visual sampling sequence of the bionic robot, and generate a visual perception region corresponding to the current control task;

[0016] Visual data of the work scene within the visual perception area is collected by a visual acquisition device set on the bionic robot body. The visual data of the work scene includes target object image data, obstacle object image data, ground area image data and operation area image data.

[0017] The motion state data of the bionic robot is synchronously collected by a state sensing device installed on the body of the bionic robot. The motion state data of the body includes the joint angle vector, posture vector, velocity vector and inertia vector of the bionic robot.

[0018] The contact status data of the foot end is collected synchronously by a contact feedback device set at the foot end of the bionic robot. The contact status data of the foot end consists of contact status marks of several foot end contact detection units. If the foot end contact detection unit is in contact state, the contact status mark is 1. If the foot end contact detection unit is in non-contact state, the contact status mark is 0.

[0019] The visual data of the work scene, the motion state data of the body, and the contact state data of the foot are aligned in time and space, and timestamps are added.

[0020] Optionally, the visual preprocessing for control tasks specifically includes:

[0021] Motion compensation processing is performed on the visual data of the work scene based on the body motion state data: an image plane motion mapping matrix is ​​constructed based on the posture data and velocity data, and coordinate transformation is performed on the visual data of the work scene based on the image plane motion mapping matrix to obtain motion-compensated visual data;

[0022] Based on the control task at the current moment, perform region enhancement processing on the motion-compensated visual data:

[0023] Based on the current control task, generate the target area mask, obstacle boundary area mask, passable area mask, and operable area mask;

[0024] A region enhancement weight map is constructed based on the target region mask, obstacle boundary region mask, passable region mask, and operable region mask;

[0025] The region augmentation weight map is multiplied element-wise with the motion-compensated visual data to obtain the region augmentation visual data.

[0026] Perform key region alignment processing on the region-enhanced visual data: extract the set of target region feature points from the region-enhanced visual data at the current and previous time steps;

[0027] Based on the set of feature points in the target region, the alignment transformation matrix of the target region is calculated by minimizing the alignment error function; according to the alignment transformation matrix of the target region, spatial mapping is performed on the region enhancement visual data of the previous time step to obtain the aligned visual data of the target region.

[0028] At each time step, the target region alignment visual data of each frame is subjected to min-max normalization to obtain standard visual control data.

[0029] Optionally, step three specifically includes:

[0030] Perform task semantic parsing based on the current control task to generate task semantic information;

[0031] Multi-scale feature extraction is performed on standard visual control data to obtain the first-scale visual feature tensor, the second-scale visual feature tensor, the third-scale visual feature tensor, and the fourth-scale visual feature tensor.

[0032] Semantic encoding of task semantic information is performed using a text Transformer encoding network structure to obtain a sequence of task semantic feature vectors;

[0033] The k-th scale visual feature tensor is convolved using a 1×1 convolution and then normalized using a Sigmoid activation function to obtain the k-th scale gated weight tensor; where k ranges from one to four.

[0034] The k-th scale gated weight tensor is multiplied element-wise with the k-th scale visual feature tensor to obtain the k-th scale gated visual feature tensor.

[0035] The motion state data of the body is encoded by linear transformation and ReLU activation function to generate a sequence of state constraint feature vectors.

[0036] The k-th scale gated visual feature tensor, the task semantic feature vector sequence, and the state constraint feature vector sequence are used to generate the k-th scale semantic enhancement feature tensor through a state constraint attention mechanism.

[0037] Based on the fourth-scale semantic enhancement feature tensor, the first-scale semantic enhancement feature tensor, the second-scale semantic enhancement feature tensor, and the third-scale semantic enhancement feature tensor are upsampled respectively, and feature concatenation and linear fusion are performed to generate the semantic enhancement feature tensor.

[0038] The semantic enhancement feature tensor and the task semantic feature vector sequence are updated through dual-granularity evolutionary query and gated temporal memory to generate a dynamic query vector set;

[0039] The dynamic query vector set, semantic enhancement feature tensor, and task semantic feature vector sequence are subjected to cross-modal decoding processing through scaling dot product attention operation and feature fusion to obtain a shared decoding feature matrix.

[0040] Based on the shared decoding feature matrix, target semantic decoding, spatial relationship decoding, and control constraint decoding are performed respectively to obtain target category vector, target position vector, target boundary vector, target relative orientation vector, obstacle constraint vector, and executable region vector, which constitute the control visual feature parameter set.

[0041] Optionally, the state-constrained attention mechanism specifically includes:

[0042] Flatten the k-th scale gated visual feature tensor in the spatial dimension and map it to the k-th scale visual query matrix through a trainable visual query mapping matrix.

[0043] The sequence of semantic feature vectors of the task is mapped to a semantic key matrix and a semantic value matrix at the k-th scale through a trainable semantic key mapping matrix and a semantic value mapping matrix, respectively.

[0044] The sequence of state constraint feature vectors is mapped to the k-th scale state constraint matrix through a trainable state constraint mapping matrix.

[0045] Add the semantic key matrix at the k-th scale to the state constraint attention term at the k-th scale to obtain the state modulation key matrix at the k-th scale. Perform a scaling dot product operation on the visual query matrix at the k-th scale and the state modulation key matrix at the k-th scale, and normalize the weights using the Softmax function to obtain the attention weight matrix at the k-th scale.

[0046] Perform matrix multiplication between the k-th scale attention weight matrix and the k-th scale semantic value matrix, and reshape the spatial dimension to obtain the k-th scale state constraint semantic feature tensor.

[0047] The semantic feature tensor of state constraint at scale k is added to the gated visual feature tensor at scale k by residual addition to obtain the semantic enhancement feature tensor at scale k.

[0048] Optionally, the dual-granularity evolution query and gated temporal memory update specifically include:

[0049] The semantic enhancement feature tensor is flattened in the spatial dimension and linearly mapped to obtain the candidate region feature matrix; the sequence of task semantic feature vectors is averaged and linearly mapped to obtain the control intent guidance vector.

[0050] The cosine similarity between the candidate region feature vector in each row of the candidate region feature matrix and the control intention guidance vector is calculated to obtain the control score of each candidate region.

[0051] All candidate regions are sorted in descending order based on the control score, and the feature vectors of the top Q candidate regions with the highest scores are selected to form a coarse-grained candidate query vector set.

[0052] The query reconstruction process is performed by crossing the coarse-grained candidate query vector set with the task semantic feature vector sequence using a cross-attention structure to obtain the semantic reconstruction query matrix;

[0053] The semantic reconstruction query matrix is ​​refined by linear transformation and ReLU activation function to obtain a set of fine-grained task query vectors.

[0054] The dynamic query vector set retained from the previous time step and the fine-grained task query vector set are concatenated along the feature dimension according to the query sequence number to obtain the memory update vector set;

[0055] The set of memory update vectors is input into the gated memory update structure, and the set of update gate coefficient vectors is generated through linear transformation and Sigmoid activation function.

[0056] The dynamic query vector set is obtained by performing element-wise weighted fusion of the fine-grained task query vector set and the query set retained in the previous time step based on the updated gate coefficient vector set.

[0057] Optionally, the control state relationship parsing and constraint information fusion specifically include:

[0058] Construct the end effector pose transformation matrix based on joint angle vectors and attitude vectors;

[0059] The position vector of the end effector in the initial coordinate system is transformed by the end effector pose transformation matrix to obtain the position vector of the end effector of the bionic robot at the current moment.

[0060] Calculate the vector difference between the target position vector and the end effector position vector to obtain the target relative position vector;

[0061] The target state vector is obtained by concatenating the target relative position vector, target boundary vector, and target relative orientation vector.

[0062] Normalize the velocity vector using the L2 norm to obtain the motion direction vector of the bionic robot at the current moment.

[0063] The conflict coefficient is obtained by calculating the cosine similarity between the obstacle constraint vector and the motion direction vector;

[0064] Based on the current control task, construct the action requirement vector; obtain the action matching coefficient by calculating the cosine similarity between the executable region vector and the action requirement vector;

[0065] Based on foot contact state data, the mean of all contact state markers is calculated to obtain the foot support coefficient;

[0066] The posture vector, inertia vector, and foot support coefficient are concatenated and linearly mapped to obtain the current posture stability vector of the bionic robot, and the target action requirement vector is constructed according to the current control task.

[0067] The attitude deviation coefficient is obtained by calculating the L2 norm of the vector difference between the attitude stability vector and the target action requirement vector.

[0068] The conflict coefficient, action matching coefficient, obstacle constraint vector, and executable region vector are concatenated to obtain the environmental constraint state vector.

[0069] The joint angle vector, posture vector, velocity vector, inertia vector, foot support coefficient, and posture deviation coefficient are concatenated to obtain the body execution motion vector.

[0070] The target state vector, environmental constraint state vector, and body execution motion vector are feature-concatenated and linearly fused to obtain the control state vector of the biomimetic robot at the current moment.

[0071] Optionally, step five specifically includes:

[0072] Based on the control state vector, the target action type corresponding to the current control task of the bionic robot is obtained by performing action type parsing through a multilayer perceptron classification structure.

[0073] Based on the environmental constraint state vector, obstacle avoidance strategy vector and actionable strategy vector are generated through a multilayer perceptron regression structure.

[0074] Based on the motion vector of the body, state parsing is performed through a gated loop unit to obtain the posture adjustment parameter vector and gait switching parameter vector;

[0075] The target action type, obstacle avoidance strategy vector, feasible action strategy vector, posture adjustment parameter vector and gait switching parameter vector are concatenated to obtain the joint motion planning vector.

[0076] The joint motion planning vectors are mapped to parameters through a multilayer perceptron regression structure to generate execution control parameters corresponding to the joint drive units, foot actuators and end effectors of the bionic robot.

[0077] The execution control parameters include joint angle control parameters, joint angular velocity control parameters, foot landing point control parameters, foot swing height control parameters, end effector position control parameters, end effector posture control parameters, and gait switching timing control parameters.

[0078] Optionally, the control commands include joint position control commands, joint speed control commands, foot trajectory control commands, end effector pose control commands, and gait switching timing control commands.

[0079] A biomimetic robot control system based on AI vision according to an embodiment of the present invention includes:

[0080] The data acquisition module is used to simultaneously collect visual data of the working scene, body motion state data, and foot contact state data of the bionic robot within the visual perception area.

[0081] The visual preprocessing module is used to perform visual preprocessing on the visual data of the operation scene in a control-oriented manner to obtain standard visual control data.

[0082] The visual recognition and feature extraction module is used to perform AI visual recognition and control visual feature extraction on standard visual control data through an improved Grounding DINO network, so as to obtain the control visual feature parameter set of the bionic robot at the current moment.

[0083] The state analysis and fusion module is used to analyze the control state relationship and fuse the constraint information of the control visual feature parameter set, the body motion state data and the foot contact state data to obtain the control state vector of the bionic robot at the current moment.

[0084] The decision planning module is used to perform hierarchical decision analysis and control planning based on the control state vector to obtain the execution control parameters of each joint drive unit, foot actuator and end effector of the bionic robot.

[0085] The control command generation and driving module is used to parse the execution control parameters into control commands through inverse kinematics solution and trajectory interpolation methods, and to drive the bionic robot based on the control commands;

[0086] The feedback correction module is used to collect execution feedback data in real time during the process of the bionic robot performing control tasks, and to perform error analysis and closed-loop control correction based on the execution feedback data.

[0087] The beneficial effects of this invention are:

[0088] This invention synchronously collects and spatiotemporally aligns visual data, body motion state data, and foot contact state data of a biomimetic robot's operational scenario. In the visual preprocessing stage, it introduces motion compensation based on body motion state, region enhancement for the current control task, and a key region temporal alignment mechanism, resulting in higher stability and control relevance of the standard visual control data entering the recognition stage. In the AI ​​visual recognition stage, an improved Grounding DINO network, incorporating state-constrained attention, dual-granularity evolutionary query, and gated temporal memory update, is used to identify, locate, and extract control visual features from target objects, obstacles, and task-related regions. This ensures the output is simultaneously constrained by body state, task semantics, and temporal continuity information, thereby improving target recognition accuracy, spatial relationship analysis capability, and control constraint representation capability. In the control state construction stage, the control visual feature parameter set is fused with body motion state data and foot contact state data through control state relationship analysis and constraint information fusion, forming a control state vector containing target state, environmental constraint state, and body execution state. This enhances the global consistency and execution targeting of subsequent hierarchical decision-making and control planning. During the execution control phase, the execution control parameters are parsed into control commands through inverse kinematics solving and trajectory interpolation methods. Error analysis and closed-loop control correction are then performed using execution feedback data, forming an integrated closed-loop control mechanism encompassing perception, analysis, planning, execution, and feedback. This reduces control deviations caused by visual jitter, state mismatch, and environmental disturbances, improving the bionic robot's environmental adaptability, motion planning accuracy, real-time control performance, and task execution stability in complex scenarios. Therefore, this invention is of great significance for improving the autonomous control level of bionic robots, reducing task failure rates, and enhancing reliable operation under complex working conditions. Attached Figure Description

[0089] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0090] Figure 1 This is a schematic diagram of a biomimetic robot control method and system based on AI vision proposed in this invention;

[0091] Figure 2 This is a flowchart of an improved Grounding DINO network structure in a biomimetic robot control method and system based on AI vision proposed in this invention.

[0092] Figure 3 This invention presents a bionic robot control method based on AI vision and a flowchart of the bionic robot control state fusion and control execution process in the system. Detailed Implementation

[0093] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0094] refer to Figures 1-3 A biomimetic robot control method based on AI vision includes the following steps:

[0095] Step 1: Simultaneously collect visual data of the working scene, motion state data of the body, and contact state data of the foot end of the bionic robot within the visual perception area;

[0096] Step 2: Perform visual preprocessing on the visual data of the work scene to obtain standard visual control data;

[0097] Step 3: Using an improved Grounding DINO network, AI visual recognition and control visual feature extraction are performed on standard visual control data to obtain the control visual feature parameter set of the bionic robot at the current moment; the improved Grounding DINO network introduces a state-constrained attention mechanism, dual-granularity evolution query, and gated temporal memory update;

[0098] Step 4: Perform control state relationship analysis and constraint information fusion on the control visual feature parameter set, body motion state data and foot contact state data to obtain the control state vector of the bionic robot at the current moment;

[0099] Step 5: Based on the control state vector, perform hierarchical decision analysis and control planning to obtain the execution control parameters of each joint drive unit, foot actuator and end effector of the bionic robot;

[0100] Step 6: The execution control parameters are parsed into control commands through inverse kinematics solution and trajectory interpolation, and the bionic robot is driven and controlled based on the control commands;

[0101] Step 7: During the process of the bionic robot performing control tasks, real-time execution feedback data is collected, and error analysis and closed-loop control correction are performed based on the execution feedback data. Specifically:

[0102] The execution feedback data includes the visual data of the current work scene, the motion state data of the body, and the contact state data of the foot end;

[0103] The visual data of the current operation scene is preprocessed for control tasks to obtain feedback standard visual control data; the feedback standard visual control data is then subjected to AI visual recognition and control feature extraction through an improved Grounding DINO network to obtain feedback control visual feature vector; the feedback control visual feature vector, the body motion state data and the foot contact state data are analyzed for control state relationship and fused with constraint information to obtain feedback control state vector.

[0104] Extract the target position vector, target boundary vector, target relative orientation vector, obstacle constraint vector, and executable region vector from the feedback control visual feature vector; extract the environmental constraint state vector and the body execution motion vector from the feedback control state vector.

[0105] The target position error is calculated based on the vector difference between the target position vector and the target reference position vector corresponding to the current control task; the control state error is calculated based on the vector difference between the feedback control state vector and the control state vector; the foot contact error is calculated based on the difference between the foot contact state data and the foot landing point control parameters; and the attitude execution error and gait switching error are calculated based on the difference between the body motion state data and the attitude adjustment parameter vector and the gait switching parameter vector, respectively.

[0106] The target position error, control state error, foot contact error, posture execution error, and gait switching error are fused to obtain the comprehensive control error; where the comprehensive control error is a normalized value.

[0107] If the overall control error is greater than or equal to the error threshold of 0.1, the vector difference between the feedback control state vector and the current control state vector is calculated to obtain the control deviation vector; the control deviation vector is linearly corrected to generate the motion planning correction vector; the motion planning correction vector and the joint motion planning vector are weighted and updated to obtain the updated joint motion planning vector; the updated joint motion planning vector is input into the multilayer perceptron regression structure to regenerate the execution control parameters.

[0108] The regenerated execution control parameters are re-analyzed into control commands through inverse kinematics solution and trajectory interpolation, and the bionic robot is then corrected for closed-loop control based on the updated control commands.

[0109] If the overall control error is less than the preset error threshold, the current control command remains unchanged, and the drive control corresponding to the current control task continues to be executed.

[0110] In this embodiment, step one specifically includes:

[0111] Obtain the visual perception center direction, visual perception range, and visual sampling sequence of the bionic robot, and generate a visual perception region corresponding to the current control task;

[0112] Visual data of the work scene within the visual perception area is collected by a visual acquisition device installed on the bionic robot body. The visual data of the work scene includes target object image data, obstacle object image data, ground area image data and operation area image data. The control tasks include grasping task, walking task, obstacle avoidance task, turning task, support adjustment task, posture adjustment task and interaction task.

[0113] The motion state data of the bionic robot is synchronously collected by a state sensing device installed on the body of the bionic robot. The motion state data of the body includes the joint angle vector, posture vector, velocity vector and inertia vector of the bionic robot. The joint angle vector is used to characterize the current rotation state of each joint module, the posture vector is used to characterize the spatial posture state of the bionic robot's torso or limbs, the velocity vector is used to characterize the motion velocity state of the bionic robot body or each moving part, and the inertia vector is used to characterize the acceleration state and angular velocity state of the bionic robot body.

[0114] The contact status data of the foot end is collected synchronously by a contact feedback device set at the foot end of the bionic robot. The contact status data of the foot end consists of contact status marks of several foot end contact detection units. If the foot end contact detection unit is in contact state, the contact status mark is 1. If the foot end contact detection unit is in non-contact state, the contact status mark is 0.

[0115] The visual data of the work scene, the motion state data of the body, and the contact state data of the foot are aligned in time and space, and timestamps are added.

[0116] In this embodiment, the visual preprocessing for control tasks specifically includes:

[0117] Motion compensation processing is performed on the visual data of the work scene based on the body motion state data: an image plane motion mapping matrix is ​​constructed based on the posture data and velocity data, and coordinate transformation is performed on the visual data of the work scene based on the image plane motion mapping matrix to obtain motion-compensated visual data;

[0118] Based on the current control task, region enhancement processing is performed on the motion-compensated visual data: According to the current control task, target region masks, obstacle boundary region masks, passable region masks, and operable region masks are generated. Specifically: The target object category, obstacle object category, passability analysis area, and operable object are determined based on the current task objective; image regions corresponding to the target object category are extracted from the motion-compensated visual data, and the pixel positions corresponding to the image region are assigned a value of 1, while the remaining positions are assigned a value of 0, generating the target region mask; image regions corresponding to the obstacle object category are extracted from the motion-compensated visual data to obtain the obstacle object region, edge detection is performed on the obstacle object region to obtain the obstacle object boundary pixel set, and the pixel positions corresponding to the obstacle object boundary pixel set are assigned a value of 1, while the remaining pixel positions are assigned a value of 0, generating the obstacle boundary region mask; ground regions or obstacle-free regions are extracted from the motion-compensated visual data, and pixel regions overlapping with the obstacle object region are removed, generating the passable region mask; image regions corresponding to the operable object are extracted from the motion-compensated visual data, and the pixel positions within the image region are assigned a value of 1, while the remaining positions are assigned a value of 0, generating the operable region mask;

[0119] Based on the target region mask, obstacle boundary region mask, passable region mask, and operable region mask, a region enhancement weight map is constructed. Specifically, the enhancement coefficients for the target region, obstacle boundary region, passable region, and operable region are set to 0.8, 0.6, 0.4, and 0.7, respectively. Element-wise multiplication operations are performed on the target region mask, obstacle boundary region mask, passable region mask, and operable region mask to obtain the target region weight component map, obstacle boundary region weight component map, passable region weight component map, and operable region weight component map. The target region weight component map, obstacle boundary region weight component map, passable region weight component map, and operable region weight component map are then added element-wise to the unit weight map to obtain the region enhancement weight map.

[0120] The region augmentation weight map is multiplied element-wise with the motion-compensated visual data to obtain the region augmentation visual data.

[0121] Perform key region alignment processing on the region-enhanced visual data: extract the set of target region feature points from the region-enhanced visual data at the current and previous time steps;

[0122] Based on the set of feature points in the target region, the alignment transformation matrix of the target region is calculated by minimizing the alignment error function. According to the alignment transformation matrix of the target region, spatial mapping is performed on the region enhancement visual data of the previous time step so that the target region at the current time step and the target region at the previous time step are in the same reference coordinate system, thereby obtaining the target region alignment visual data.

[0123] At each time step, the target region alignment visual data of each frame is subjected to min-max normalization to obtain standard visual control data.

[0124] In this embodiment, step three specifically includes:

[0125] Perform task semantic parsing based on the current control task to generate task semantic information;

[0126] Multi-scale feature extraction is performed on standard visual control data to obtain the first-scale visual feature tensor, the second-scale visual feature tensor, the third-scale visual feature tensor, and the fourth-scale visual feature tensor.

[0127] Semantic encoding of task semantic information is performed using a text Transformer encoding network structure to obtain a sequence of task semantic feature vectors;

[0128] The k-th scale visual feature tensor is convolved using a 1×1 convolution and then normalized using a Sigmoid activation function to obtain the k-th scale gated weight tensor; where k ranges from one to four.

[0129] The k-th scale gated weight tensor is multiplied element-wise with the k-th scale visual feature tensor to obtain the k-th scale gated visual feature tensor.

[0130] The motion state data of the body is encoded by linear transformation and ReLU activation function to generate a sequence of state constraint feature vectors.

[0131] The k-th scale gated visual feature tensor, the task semantic feature vector sequence, and the state constraint feature vector sequence are used to generate the k-th scale semantic enhancement feature tensor through a state constraint attention mechanism.

[0132] Based on the fourth-scale semantic enhancement feature tensor, the first-scale semantic enhancement feature tensor, the second-scale semantic enhancement feature tensor, and the third-scale semantic enhancement feature tensor are upsampled respectively, and feature concatenation and linear fusion are performed to generate the semantic enhancement feature tensor.

[0133] The semantic enhancement feature tensor and the task semantic feature vector sequence are updated through dual-granularity evolutionary query and gated temporal memory to generate a dynamic query vector set;

[0134] The dynamic query vector set, semantic enhancement feature tensor, and task semantic feature vector sequence are subjected to cross-modal decoding processing through scaling dot product attention operation and feature fusion to obtain a shared decoding feature matrix.

[0135] Based on the shared decoding feature matrix, target semantic decoding, spatial relationship decoding, and control constraint decoding are performed respectively to obtain target category vector, target position vector, target boundary vector, target relative orientation vector, obstacle constraint vector, and executable region vector, which constitute the control visual feature parameter set.

[0136] In this embodiment, the state-constrained attention mechanism specifically includes:

[0137] Flatten the k-th scale gated visual feature tensor in the spatial dimension and map it to the k-th scale visual query matrix through a trainable visual query mapping matrix.

[0138] The sequence of semantic feature vectors of the task is mapped to a semantic key matrix and a semantic value matrix at the k-th scale through a trainable semantic key mapping matrix and a semantic value mapping matrix, respectively.

[0139] The sequence of state constraint feature vectors is mapped to the k-th scale state constraint matrix through a trainable state constraint mapping matrix.

[0140] Add the semantic key matrix at the k-th scale to the state constraint attention term at the k-th scale to obtain the state modulation key matrix at the k-th scale. Perform a scaling dot product operation on the visual query matrix at the k-th scale and the state modulation key matrix at the k-th scale, and normalize the weights using the Softmax function to obtain the attention weight matrix at the k-th scale.

[0141] Perform matrix multiplication between the k-th scale attention weight matrix and the k-th scale semantic value matrix, and reshape the spatial dimension to obtain the k-th scale state constraint semantic feature tensor.

[0142] The semantic feature tensor of state constraint at scale k is added to the gated visual feature tensor at scale k by residual addition to obtain the semantic enhancement feature tensor at scale k.

[0143] In this embodiment, the dual-granularity evolution query and gated temporal memory update specifically include:

[0144] The semantic enhancement feature tensor is flattened in the spatial dimension and linearly mapped to obtain the candidate region feature matrix; the sequence of task semantic feature vectors is averaged and linearly mapped to obtain the control intent guidance vector.

[0145] The cosine similarity between the candidate region feature vector in each row of the candidate region feature matrix and the control intention guidance vector is calculated to obtain the control score of each candidate region.

[0146] All candidate regions are sorted in descending order based on the control score, and the feature vectors of the top Q candidate regions with the highest scores are selected to form a coarse-grained candidate query vector set.

[0147] The query reconstruction process is performed by crossing the coarse-grained candidate query vector set with the task semantic feature vector sequence using a cross-attention structure to obtain the semantic reconstruction query matrix;

[0148] The semantic reconstruction query matrix is ​​refined by linear transformation and ReLU activation function to obtain a set of fine-grained task query vectors.

[0149] The dynamic query vector set retained from the previous time step and the fine-grained task query vector set are concatenated along the feature dimension according to the query sequence number to obtain the memory update vector set;

[0150] The set of memory update vectors is input into the gated memory update structure, and the set of update gate coefficient vectors is generated through linear transformation and Sigmoid activation function.

[0151] The dynamic query vector set is obtained by performing element-wise weighted fusion of the fine-grained task query vector set and the query set retained in the previous time step based on the updated gate coefficient vector set.

[0152] In this invention, the improved Grounding DINO network maintains the same overall architecture as the standard Grounding DINO network. Both belong to the category of joint visual-language object detection networks, and both focus on visual feature extraction, text semantic encoding, cross-modal feature interaction, query generation, and decoding output. The standard Grounding DINO network typically includes an image encoding backbone, a text encoding backbone, a feature enhancement module, a language-guided query selection module, and a cross-modal decoding module. The image encoding backbone extracts multi-scale visual features from the input image, while the text encoding backbone extracts semantic features from the text prompts. Cross-modal interaction establishes the correspondence between visual and semantic features, and the query selection and decoding process outputs the target category and target location results. Therefore, the improved Grounding DINO network does not change the basic technical approach of the standard Grounding DINO network.

[0153] The improved Grounding DINO network enhances the intermediate key modules for biomimetic robot control tasks based on the standard Grounding DINO network. In the visual and semantic feature interaction stage, a state-constrained attention mechanism is introduced, allowing the ontological motion state data to participate in attention weight modulation after state encoding, thereby adding state constraint information during visual-semantic fusion. In the query generation stage, a dual-granularity evolutionary query and gated temporal memory update structure are introduced. First, a coarse-grained candidate query completes the region selection, then a fine-grained query reconstructs the task-related query, and the temporal memory is updated by combining the dynamic query from the previous time step. In the decoding output stage, the shared decoding feature matrix is ​​further decomposed into target semantic decoding, spatial relationship decoding, and control constraint decoding to output a set of control visual feature parameters suitable for the control task.

[0154] By introducing a state-constrained attention mechanism, the improved Grounding DINO network can combine the ontological motion state to suppress feature shifts caused by posture changes, body vibrations, and motion disturbances during visual recognition, thus improving recognition stability. Through dual-granularity evolutionary query and gated temporal memory update, the network can enhance its ability to continuously track task-related targets, reducing query drift and target loss in dynamic scenes. By decomposing and decoding target semantics, spatial relationships, and control constraints, the network output is no longer limited to target categories and bounding boxes, improving the accuracy and real-time performance of subsequent control state construction, hierarchical decision planning, and execution control.

[0155] In this embodiment, the control state relationship parsing and constraint information fusion specifically include:

[0156] Construct the end effector pose transformation matrix based on joint angle vectors and attitude vectors;

[0157] The position vector of the end effector in the initial coordinate system is transformed by the end effector pose transformation matrix to obtain the position vector of the end effector of the bionic robot at the current moment.

[0158] Calculate the vector difference between the target position vector and the end effector position vector to obtain the target relative position vector;

[0159] The target state vector is obtained by concatenating the target relative position vector, target boundary vector, and target relative orientation vector.

[0160] Normalize the velocity vector using the L2 norm to obtain the motion direction vector of the bionic robot at the current moment.

[0161] The conflict coefficient is obtained by calculating the cosine similarity between the obstacle constraint vector and the motion direction vector;

[0162] Construct an action requirement vector based on the current control task;

[0163] The action matching coefficient is obtained by calculating the cosine similarity between the executable region vector and the action requirement vector;

[0164] Based on foot contact state data, the mean of all contact state markers is calculated to obtain the foot support coefficient;

[0165] The posture vector, inertia vector, and foot support coefficient are concatenated and linearly mapped to obtain the current posture stability vector of the bionic robot, and the target action requirement vector is constructed according to the current control task.

[0166] The attitude deviation coefficient is obtained by calculating the L2 norm of the vector difference between the attitude stability vector and the target action requirement vector.

[0167] The conflict coefficient, action matching coefficient, obstacle constraint vector, and executable region vector are concatenated to obtain the environmental constraint state vector.

[0168] The joint angle vector, posture vector, velocity vector, inertia vector, foot support coefficient, and posture deviation coefficient are concatenated to obtain the body execution motion vector.

[0169] The target state vector, environmental constraint state vector, and body execution motion vector are feature-concatenated and linearly fused to obtain the control state vector of the biomimetic robot at the current moment.

[0170] In this invention, the conflict coefficient is used to quantify the conflict relationship between the obstacle area and the current motion direction of the bionic robot, the action matching coefficient is used to quantify the matching relationship between the executable area and the current control task, and the posture deviation coefficient is used to quantify the deviation relationship between the current posture stability of the bionic robot and the target action requirements.

[0171] In this embodiment, step five specifically includes:

[0172] Based on the control state vector, the target action type corresponding to the current control task of the bionic robot is obtained by performing action type parsing through a multilayer perceptron classification structure.

[0173] Based on the environmental constraint state vector, obstacle avoidance strategy vector and actionable strategy vector are generated through a multilayer perceptron regression structure.

[0174] Based on the motion vector of the body, state parsing is performed through a gated loop unit to obtain the posture adjustment parameter vector and gait switching parameter vector;

[0175] The target action type, obstacle avoidance strategy vector, feasible action strategy vector, posture adjustment parameter vector and gait switching parameter vector are concatenated to obtain the joint motion planning vector.

[0176] The joint motion planning vectors are mapped to parameters through a multilayer perceptron regression structure to generate execution control parameters corresponding to the joint drive units, foot actuators and end effectors of the bionic robot.

[0177] The execution control parameters include joint angle control parameters, joint angular velocity control parameters, foot landing point control parameters, foot swing height control parameters, end effector position control parameters, end effector posture control parameters, and gait switching timing control parameters.

[0178] In this embodiment, the control commands include joint position control commands, joint speed control commands, foot trajectory control commands, end effector pose control commands, and gait switching timing control commands.

[0179] A biomimetic robot control system based on AI vision includes:

[0180] The data acquisition module is used to simultaneously collect visual data of the working scene, body motion state data, and foot contact state data of the bionic robot within the visual perception area.

[0181] The visual preprocessing module is used to perform visual preprocessing on the visual data of the operation scene in a control-oriented manner to obtain standard visual control data.

[0182] The visual recognition and feature extraction module is used to perform AI visual recognition and control visual feature extraction on standard visual control data through an improved Grounding DINO network, so as to obtain the control visual feature parameter set of the bionic robot at the current moment.

[0183] The state analysis and fusion module is used to analyze the control state relationship and fuse the constraint information of the control visual feature parameter set, the body motion state data and the foot contact state data to obtain the control state vector of the bionic robot at the current moment.

[0184] The decision planning module is used to perform hierarchical decision analysis and control planning based on the control state vector to obtain the execution control parameters of each joint drive unit, foot actuator and end effector of the bionic robot.

[0185] The control command generation and driving module is used to parse the execution control parameters into control commands through inverse kinematics solution and trajectory interpolation methods, and to drive the bionic robot based on the control commands;

[0186] The feedback correction module is used to collect execution feedback data in real time during the process of the bionic robot performing control tasks, and to perform error analysis and closed-loop control correction based on the execution feedback data.

[0187] Example 1: To verify the feasibility of this invention in practice, the method was applied to a complex indoor warehouse sorting and cross-regional handling scenario. The work area was 30m × 20m, with the ground including an epoxy flooring area, a rubber buffer area, and a slightly sloping transition area. Shelves, turnover boxes, pallets, temporary obstacles, ground warning lines, and workbenches were arranged within the area. The bionic robot adopted a quadrupedal structure with a single arm. The robot body was 1.15m high and weighed 42kg. It was equipped with a binocular vision acquisition device, an inertial measurement unit, a joint angle encoder, a foot contact detection unit, and an end effector pose sensing unit. The vision acquisition device had an output resolution of 1920×1080 and a sampling frequency of 30fps; the sampling frequency for joint angle, body posture, velocity, and inertial data was 200Hz; and the sampling frequency for foot contact status data was 500Hz. The robot's specific tasks are: to identify designated turnover boxes during dynamic walking, grab them from the front of the shelf and transport them to the operating table area, autonomously avoid randomly appearing pedestrian models, moving carts and scattered obstacles on the ground, and maintain posture stability and foot support safety in areas where different ground materials are switched.

[0188] In the implementation of this invention, a vision acquisition device installed on the head of the bionic robot continuously acquires images of the work scene in front, and dynamically determines the visual perception area by combining the robot's current movement direction and task mode. The state sensing device inside the robot body simultaneously acquires joint angle vectors, posture vectors, velocity vectors, and inertial vectors, while the foot contact detection unit simultaneously outputs contact status markers for the four feet. All data is time- and space-aligned in the controller and a unified timestamp is added. Subsequently, visual preprocessing for control tasks is performed on the visual data of the work scene to obtain standard visual control data.

[0189] In the visual recognition and control feature extraction stage, standard visual control data is used to generate a control visual feature parameter set through an improved Grounding DINO network. This control visual feature parameter set is then further analyzed and constraint information is fused with the body motion state data and foot contact state data. The controller first constructs the end effector pose transformation relationship based on joint angle vectors and posture vectors, calculates the relative position state between the target turnover box and the end effector, and then combines the velocity vector to obtain the current motion direction. The relationship between obstacle constraint vectors and motion direction is used to generate a conflict degree representation. Simultaneously, combining the executable area vector and the current action requirement, the matching degree between the next landing area and the grasping action is determined, and the posture vector, inertia vector, and foot support state are jointly mapped to posture stability information. The control state vector constructed in this way not only describes "what is seen" but also "whether the robot can execute safely and how it should prioritize execution," thus providing sufficient control semantic support for subsequent hierarchical decision-making.

[0190] In the decision-making and planning stage, this invention completes action type parsing, obstacle avoidance strategy generation, posture adjustment, and gait switching control based on the control state vector. When the robot is far from the target, it primarily uses a gait with high throughput. When a conflict is detected between the robot and the current direction of travel, the system prioritizes outputting an obstacle avoidance lateral movement strategy. When approaching the target cargo box, the robot switches to a low-speed, precise positioning mode, while adjusting the robot's height and end effector posture to ensure a smooth grasping trajectory. When the foot contact status data shows insufficient support for a particular forefoot, the robot reduces its stride frequency and adjusts its foot placement to ensure coordination between the grasping action and overall stability. All execution control parameters are further analyzed into joint position control commands, joint velocity control commands, foot trajectory control commands, end effector pose control commands, and gait switching timing control commands through inverse kinematics and trajectory interpolation methods, driving the bionic robot to complete the entire handling task. During execution, the system continuously collects new visual data, body state data, and foot contact feedback data. When it detects a shift in the target position, a decrease in posture stability, or an error in the path, it immediately performs error analysis and closed-loop control correction to maintain task continuity and execution stability.

[0191] To further verify the practical effects of this invention, three comparative schemes were set up. Comparative scheme A is a method based on traditional image detection and rule control, which uses a conventional object detection network to output target boxes and uses preset threshold rules to complete obstacle avoidance and grasping control. Comparative scheme B is a method based on standard Grounding DINO and conventional state fusion. In the visual recognition stage, it uses an unmodified Grounding DINO network, but does not include state-constrained attention mechanism, dual-granularity evolutionary query and gated temporal memory update, nor does it use the control-oriented visual preprocessing proposed in this invention. The experiment ran continuously for 20 working days, with 50 complete handling tasks executed each day, for a total of 1000 task samples. The experimental results are shown in Table 1.

[0192] Table 1. Performance Comparison of Different Methods in Complex Warehouse Handling Scenarios

[0193] Target recognition accuracy / % 89.6 93.8 97.5 Obstacle recognition accuracy / % 86.9 92.4 96.8 Target relative position error / cm 6.8 4.2 2.1 Average number of path replanning attempts 3.7 2.5 1.3 Average task completion time / s 41.5 35.8 29.6 Foot slip rate / % 7.4 4.6 1.8 Crawling success rate / % 88.1 93.2 97.1 Overall task success rate / % 81.7 89.4 95.8 Closed-loop correction response time / ms 132 96 61 Average control delay / ms 98 74 49

[0194] As shown in Table 1, the method of this invention outperforms both comparative scheme A and comparative scheme B in all performance indicators under complex warehousing and handling scenarios. In terms of perception performance, the target recognition accuracy of the method of this invention reaches 97.5%, and the obstacle recognition accuracy reaches 96.8%, significantly higher than 89.6% and 86.9% of comparative scheme A, and higher than 93.8% and 92.4% of comparative scheme B. This indicates that after visual preprocessing oriented towards control tasks and AI visual recognition and control feature extraction using an improved Grounding DINO network, this invention can more accurately identify target objects and obstacle areas, thereby improving the reliability of visual perception in complex scenarios.

[0195] In terms of control accuracy, the target relative position error of the method of this invention is only 2.1 cm, significantly lower than 6.8 cm for comparative scheme A and 4.2 cm for comparative scheme B, indicating that the present invention has higher accuracy in target positioning, spatial relationship analysis, and control state construction. Meanwhile, the average number of path replanning iterations of the method of this invention is only 1.3 times, lower than 3.7 times for comparative scheme A and 2.5 times for comparative scheme B, demonstrating that the present invention can more accurately complete environmental constraint analysis and initial action planning, reducing repeated corrections during execution.

[0196] In terms of execution stability, the method of this invention exhibits a foot slippage rate of 1.8%, a grasping success rate of 97.1%, and an overall task success rate of 95.8%, all of which are superior to the two comparative schemes. This demonstrates that the present invention has significant advantages in foot support stability control, end effector target approach, and overall task coordinated execution. Furthermore, the average task completion time of the method of this invention is reduced to 29.6s, the closed-loop correction response time is reduced to 61ms, and the average control latency is reduced to 49ms, indicating that the present invention not only improves control accuracy and execution stability but also possesses good real-time response capabilities.

[0197] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A biomimetic robot control method based on AI vision, characterized in that, Includes the following steps: Step 1: Simultaneously collect visual data of the working scene, motion state data of the body, and contact state data of the foot end of the bionic robot within the visual perception area; Step 2: Perform visual preprocessing on the visual data of the work scene to obtain standard visual control data; Step 3: Using an improved GroundingDINO network, AI visual recognition and control visual feature extraction are performed on standard visual control data to obtain the control visual feature parameter set of the bionic robot at the current moment; the improved GroundingDINO network introduces a state-constrained attention mechanism, dual-granularity evolution query, and gated temporal memory update. Step 4: Perform control state relationship analysis and constraint information fusion on the control visual feature parameter set, body motion state data and foot contact state data to obtain the control state vector of the bionic robot at the current moment; Step 5: Based on the control state vector, perform hierarchical decision analysis and control planning to obtain the execution control parameters of each joint drive unit, foot actuator and end effector of the bionic robot; Step 6: The execution control parameters are parsed into control commands through inverse kinematics solution and trajectory interpolation, and the bionic robot is driven and controlled based on the control commands; Step 7: During the process of the bionic robot performing control tasks, collect execution feedback data in real time, and perform error analysis and closed-loop control correction based on the execution feedback data.

2. The biomimetic robot control method based on AI vision according to claim 1, characterized in that, Step one specifically includes: Obtain the visual perception center direction, visual perception range, and visual sampling sequence of the bionic robot, and generate a visual perception region corresponding to the current control task; Visual data of the work scene within the visual perception area is collected by a visual acquisition device set on the bionic robot body. The visual data of the work scene includes target object image data, obstacle object image data, ground area image data and operation area image data. The motion state data of the bionic robot is synchronously collected by a state sensing device installed on the body of the bionic robot. The motion state data of the body includes the joint angle vector, posture vector, velocity vector and inertia vector of the bionic robot. The contact status data of the foot end is collected synchronously by a contact feedback device set at the foot end of the bionic robot. The contact status data of the foot end consists of contact status marks of several foot end contact detection units. If the foot end contact detection unit is in contact state, the contact status mark is 1. If the foot end contact detection unit is in non-contact state, the contact status mark is 0. The visual data of the work scene, the motion state data of the body, and the contact state data of the foot are aligned in time and space, and timestamps are added.

3. The biomimetic robot control method based on AI vision according to claim 1, characterized in that, The visual preprocessing for control tasks specifically includes: Motion compensation processing is performed on the visual data of the work scene based on the body motion state data: an image plane motion mapping matrix is ​​constructed based on the posture data and velocity data, and coordinate transformation is performed on the visual data of the work scene based on the image plane motion mapping matrix to obtain motion-compensated visual data; Based on the control task at the current moment, perform region enhancement processing on the motion-compensated visual data: Based on the current control task, generate the target area mask, obstacle boundary area mask, passable area mask, and operable area mask; A region enhancement weight map is constructed based on the target region mask, obstacle boundary region mask, passable region mask, and operable region mask; The region augmentation weight map is multiplied element-wise with the motion-compensated visual data to obtain the region augmentation visual data. Perform key region alignment processing on the region-enhanced visual data: extract the set of target region feature points from the region-enhanced visual data at the current and previous time steps; Based on the set of feature points in the target region, the alignment transformation matrix of the target region is calculated by minimizing the alignment error function; according to the alignment transformation matrix of the target region, spatial mapping is performed on the region enhancement visual data of the previous time step to obtain the aligned visual data of the target region. At each time step, the target region alignment visual data of each frame is subjected to min-max normalization to obtain standard visual control data.

4. The biomimetic robot control method based on AI vision according to claim 1, characterized in that, Step three specifically includes: Perform task semantic parsing based on the current control task to generate task semantic information; Multi-scale feature extraction is performed on standard visual control data to obtain the first-scale visual feature tensor, the second-scale visual feature tensor, the third-scale visual feature tensor, and the fourth-scale visual feature tensor. Semantic encoding of task semantic information is performed using a text Transformer encoding network structure to obtain a sequence of task semantic feature vectors; The k-th scale visual feature tensor is convolved using a 1×1 convolution and then normalized using a Sigmoid activation function to obtain the k-th scale gated weight tensor; where k ranges from one to four. The k-th scale gated weight tensor is multiplied element-wise with the k-th scale visual feature tensor to obtain the k-th scale gated visual feature tensor. The motion state data of the body is encoded by linear transformation and ReLU activation function to generate a sequence of state constraint feature vectors. The k-th scale gated visual feature tensor, the task semantic feature vector sequence, and the state constraint feature vector sequence are used to generate the k-th scale semantic enhancement feature tensor through a state constraint attention mechanism. Based on the fourth-scale semantic enhancement feature tensor, the first-scale semantic enhancement feature tensor, the second-scale semantic enhancement feature tensor, and the third-scale semantic enhancement feature tensor are upsampled respectively, and feature concatenation and linear fusion are performed to generate the semantic enhancement feature tensor. The semantic enhancement feature tensor and the task semantic feature vector sequence are updated through dual-granularity evolutionary query and gated temporal memory to generate a dynamic query vector set; The dynamic query vector set, semantic enhancement feature tensor, and task semantic feature vector sequence are subjected to cross-modal decoding processing through scaling dot product attention operation and feature fusion to obtain a shared decoding feature matrix. Based on the shared decoding feature matrix, target semantic decoding, spatial relationship decoding, and control constraint decoding are performed respectively to obtain target category vector, target position vector, target boundary vector, target relative orientation vector, obstacle constraint vector, and executable region vector, which constitute the control visual feature parameter set.

5. The biomimetic robot control method based on AI vision according to claim 1, characterized in that, The state-constrained attention mechanism specifically includes: Flatten the k-th scale gated visual feature tensor in the spatial dimension and map it to the k-th scale visual query matrix through a trainable visual query mapping matrix. The sequence of semantic feature vectors of the task is mapped to a semantic key matrix and a semantic value matrix at the k-th scale through a trainable semantic key mapping matrix and a semantic value mapping matrix, respectively. The sequence of state constraint feature vectors is mapped to the k-th scale state constraint matrix through a trainable state constraint mapping matrix. Add the semantic key matrix at the k-th scale to the state constraint attention term at the k-th scale to obtain the state modulation key matrix at the k-th scale. Perform a scaling dot product operation on the visual query matrix at the k-th scale and the state modulation key matrix at the k-th scale, and normalize the weights using the Softmax function to obtain the attention weight matrix at the k-th scale. Perform matrix multiplication between the k-th scale attention weight matrix and the k-th scale semantic value matrix, and reshape the spatial dimension to obtain the k-th scale state constraint semantic feature tensor. The semantic feature tensor of state constraint at scale k is added to the gated visual feature tensor at scale k by residual addition to obtain the semantic enhancement feature tensor at scale k.

6. The biomimetic robot control method based on AI vision according to claim 1, characterized in that, The dual-granularity evolution query and gated temporal memory update specifically include: The semantic enhancement feature tensor is flattened in the spatial dimension and linearly mapped to obtain the candidate region feature matrix; the sequence of task semantic feature vectors is averaged and linearly mapped to obtain the control intent guidance vector. The cosine similarity between the candidate region feature vector in each row of the candidate region feature matrix and the control intention guidance vector is calculated to obtain the control score of each candidate region. All candidate regions are sorted in descending order based on the control score, and the feature vectors of the top Q candidate regions with the highest scores are selected to form a coarse-grained candidate query vector set. The query reconstruction process is performed by crossing the coarse-grained candidate query vector set with the task semantic feature vector sequence using a cross-attention structure to obtain the semantic reconstruction query matrix; The semantic reconstruction query matrix is ​​refined by linear transformation and ReLU activation function to obtain a set of fine-grained task query vectors. The dynamic query vector set retained from the previous time step and the fine-grained task query vector set are concatenated along the feature dimension according to the query sequence number to obtain the memory update vector set; The set of memory update vectors is input into the gated memory update structure, and the set of update gate coefficient vectors is generated through linear transformation and Sigmoid activation function. The dynamic query vector set is obtained by performing element-wise weighted fusion of the fine-grained task query vector set and the query set retained in the previous time step based on the updated gate coefficient vector set.

7. The biomimetic robot control method based on AI vision according to claim 1, characterized in that, The control state relationship analysis and constraint information fusion specifically include: Construct the end effector pose transformation matrix based on joint angle vectors and attitude vectors; The position vector of the end effector in the initial coordinate system is transformed by the end effector pose transformation matrix to obtain the position vector of the end effector of the bionic robot at the current moment. Calculate the vector difference between the target position vector and the end effector position vector to obtain the target relative position vector; The target state vector is obtained by concatenating the target relative position vector, target boundary vector, and target relative orientation vector. Normalize the velocity vector using the L2 norm to obtain the motion direction vector of the bionic robot at the current moment. The conflict coefficient is obtained by calculating the cosine similarity between the obstacle constraint vector and the motion direction vector; Based on the current control task, construct the action requirement vector; obtain the action matching coefficient by calculating the cosine similarity between the executable region vector and the action requirement vector; Based on foot contact state data, the mean of all contact state markers is calculated to obtain the foot support coefficient; The posture vector, inertia vector, and foot support coefficient are concatenated and linearly mapped to obtain the current posture stability vector of the bionic robot, and the target action requirement vector is constructed according to the current control task. The attitude deviation coefficient is obtained by calculating the L2 norm of the vector difference between the attitude stability vector and the target action requirement vector. The conflict coefficient, action matching coefficient, obstacle constraint vector, and executable region vector are concatenated to obtain the environmental constraint state vector. The joint angle vector, posture vector, velocity vector, inertia vector, foot support coefficient, and posture deviation coefficient are concatenated to obtain the body execution motion vector. The target state vector, environmental constraint state vector, and body execution motion vector are feature-concatenated and linearly fused to obtain the control state vector of the biomimetic robot at the current moment.

8. The biomimetic robot control method based on AI vision according to claim 1, characterized in that, Step five specifically includes: Based on the control state vector, the target action type corresponding to the current control task of the bionic robot is obtained by performing action type parsing through a multilayer perceptron classification structure. Based on the environmental constraint state vector, obstacle avoidance strategy vector and actionable strategy vector are generated through a multilayer perceptron regression structure. Based on the motion vector of the body, state parsing is performed through a gated loop unit to obtain the posture adjustment parameter vector and gait switching parameter vector; The target action type, obstacle avoidance strategy vector, feasible action strategy vector, posture adjustment parameter vector and gait switching parameter vector are concatenated to obtain the joint motion planning vector. The joint motion planning vectors are mapped to parameters through a multilayer perceptron regression structure to generate execution control parameters corresponding to the joint drive units, foot actuators and end effectors of the bionic robot. The execution control parameters include joint angle control parameters, joint angular velocity control parameters, foot landing point control parameters, foot swing height control parameters, end effector position control parameters, end effector posture control parameters, and gait switching timing control parameters.

9. The bionic robot control method based on AI vision according to claim 1, characterized in that, The control commands include joint position control commands, joint velocity control commands, foot trajectory control commands, end effector pose control commands, and gait switching timing control commands.

10. A bionic robot control system based on AI vision, executing the bionic robot control method based on AI vision as described in any one of claims 1 to 9, characterized in that, include: The data acquisition module is used to simultaneously collect visual data of the working scene, body motion state data, and foot contact state data of the bionic robot within the visual perception area. The visual preprocessing module is used to perform visual preprocessing on the visual data of the operation scene in a control-oriented manner to obtain standard visual control data. The visual recognition and feature extraction module is used to perform AI visual recognition and control visual feature extraction on standard visual control data through an improved Grounding DINO network, so as to obtain the control visual feature parameter set of the bionic robot at the current moment. The state analysis and fusion module is used to analyze the control state relationship and fuse the constraint information of the control visual feature parameter set, the body motion state data and the foot contact state data to obtain the control state vector of the bionic robot at the current moment. The decision planning module is used to perform hierarchical decision analysis and control planning based on the control state vector to obtain the execution control parameters of each joint drive unit, foot actuator and end effector of the bionic robot. The control command generation and driving module is used to parse the execution control parameters into control commands through inverse kinematics solution and trajectory interpolation methods, and to drive the bionic robot based on the control commands; The feedback correction module is used to collect execution feedback data in real time during the process of the bionic robot performing control tasks, and to perform error analysis and closed-loop control correction based on the execution feedback data.