Humanoid robot remote control method, device and equipment based on wireless communication
Through multimodal semantic parsing and reverse causal reasoning technology, combined with two-way expectation calibration and paradox verification, a complete remote control system was constructed, which solved the problems of insufficient control intention recognition and imperfect safety protection in existing technologies, and achieved high-precision, safe and reliable remote control.
Patent Information
- Application Number
- CN202511122431.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-08-12
AI Technical Summary
Existing remote control methods lack intelligent semantic understanding and intent recognition, the control sequence planning does not have reverse reasoning and security verification capabilities, the safety protection mechanism is imperfect, and the execution monitoring system is not real-time, resulting in insufficient control accuracy and safety reliability.
By integrating multimodal semantic analysis and reverse causal reasoning technology, a two-way expectation calibration channel is established. Through paradox verification and interlocking safety control, a complete remote control system is constructed to achieve intelligent conversion of control intentions and multiple safety protections.
It improves control accuracy and safety and reliability, can monitor in real time and predictively adjust control strategies, prevent potential risks, and ensure stable control of the robot in complex scenarios.
Smart Images

Figure CN120620231A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robot control technology, and in particular to a method, device and equipment for remotely controlling a humanoid robot based on wireless communication. Background Art
[0002] Humanoid robots, a key development direction in intelligent manufacturing and service robotics, demonstrate enormous potential for applications in areas such as elderly care, medical rehabilitation, and home services. With the accelerated aging of the population and the rapidly growing demand for remote services, wireless communication-based remote control of humanoid robots has become a key enabling technology for intelligent nursing services and refined assisted operations. Remote control technology transcends geographic limitations, enabling professional operators to remotely guide robots to complete complex nursing tasks, providing an effective solution for addressing the uneven allocation of nursing resources and the shortage of qualified personnel.
[0003] However, existing remote control methods have significant technical limitations: Control command processing lacks intelligent semantic understanding and intent recognition mechanisms, making it impossible to accurately grasp the operator's true needs and desired goals; control sequence planning often uses forward planning methods, lacking reverse reasoning and safety verification capabilities, making it difficult to effectively avoid potential risks while ensuring execution effectiveness; safety protection mechanisms are simple and crude, lacking multiple interlock verifications and dynamic boundary detection, which are prone to security vulnerabilities and control failures; and the execution monitoring system is imperfect, lacking real-time state sampling and convergence analysis capabilities, making it impossible to detect deviations and adjust strategies in a timely manner. These problems seriously restrict the control accuracy, safety, reliability, and intelligence level of humanoid robots in complex scenarios. Summary of the Invention
[0004] This invention provides a method, device, and equipment for remotely controlling a humanoid robot based on wireless communication, aiming to address key technical challenges faced by existing remote control technologies in terms of control understanding, safety verification, and execution monitoring. This technical solution integrates multimodal semantic parsing and reverse causal reasoning techniques to achieve intelligent conversion from control instructions to execution sequences. It also incorporates paradox verification analysis and interlocking safety control techniques to implement multiple safety protection mechanisms. Furthermore, it incorporates predictive feedback correction technology to achieve dynamic optimization of control strategies, thereby constructing a technical system with comprehensive intelligent remote control capabilities.
[0005] A first aspect of the present invention provides a method for remotely controlling a humanoid robot based on wireless communication, comprising the following steps: receiving a remote control command from an operator, performing semantic analysis on the remote control command to generate control intention data, establishing a two-way expectation calibration channel based on the control intention data, and generating an expected consistency parameter and a deviation tolerance threshold based on the two-way expectation calibration channel; Obtaining an expected execution result based on the expected consistency parameter, inferring an optimal control sequence based on the expected execution result, and performing negative space analysis on the optimal control sequence to identify a prohibited action set; Combining the optimal control sequence and the prohibited action set to construct a paradox verification scenario, and performing a conflict test on the paradox verification scenario to generate control boundary parameters; establishing an interlock verification condition based on the manipulation boundary parameter and the deviation tolerance threshold, and generating a safe unlocking sequence based on the interlock verification condition; Parallel monitoring of the optimal control sequence based on the security unlocking sequence to obtain execution verification data, determining a target execution branch according to the execution verification data, and generating a unified instruction stream based on the target execution branch; The unified instruction stream is transmitted to the robot side via wireless communication, the robot side performs instruction verification based on the safety unlocking sequence to generate a verification pass signal, and the robot is driven to execute and obtain actual execution status data based on the verification pass signal; A convergence analysis is performed based on the actual execution state data and the deviation tolerance threshold to generate a convergence monitoring result, a causal feedback chain is obtained according to the convergence monitoring result, and the causal feedback chain is used to predictively correct subsequent control behaviors to complete remote control.
[0006] A second aspect of the present invention provides a remote control device for a humanoid robot based on wireless communication, comprising: a command parsing module, configured to receive a remote control command from an operator, perform semantic parsing on the remote control command to generate control intention data, establish a two-way desired calibration channel based on the control intention data, and generate desired consistency parameters and a deviation tolerance threshold based on the two-way desired calibration channel; a reverse reasoning module, configured to obtain an expected execution result based on the expected consistency parameter, reversely infer an optimal control sequence based on the expected execution result, and perform negative space analysis on the optimal control sequence to identify a set of prohibited actions; A paradox verification module, configured to construct a paradox verification scenario by combining the optimal control sequence and the prohibited action set, and to generate control boundary parameters by performing a conflict test on the paradox verification scenario; a safety control module, configured to establish an interlock verification condition based on the manipulation boundary parameter and the deviation tolerance threshold, and generate a safety unlocking sequence based on the interlock verification condition; an execution monitoring module, configured to monitor the optimal control sequence in parallel based on the security unlocking sequence to obtain execution verification data, determine a target execution branch based on the execution verification data, and generate a unified instruction stream based on the target execution branch; a communication execution module, configured to transmit the unified instruction stream to the robot side via wireless communication, perform instruction verification on the robot side based on the safety unlock sequence to generate a verification pass signal, and drive the robot to execute and obtain actual execution status data based on the verification pass signal; A feedback optimization module is used to perform convergence analysis based on the actual execution state data and the deviation tolerance threshold to generate a convergence monitoring result, obtain a causal feedback chain based on the convergence monitoring result, and use the causal feedback chain to predictively correct subsequent control behaviors to complete remote control.
[0007] The third aspect of the present invention proposes a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of a method for remotely controlling a humanoid robot based on wireless communication disclosed in the first aspect are implemented.
[0008] The beneficial effects of this invention are reflected in the following aspects: First, through multimodal semantic parsing technology, it can simultaneously process multiple input signals such as voice, gestures, and EEG, more accurately understanding the operator's true intentions. A two-way expectation calibration mechanism aligns control requirements with the robot's actual capabilities, avoiding the problems of "overly high expectations leading to inability to execute" or "wasted capabilities." Using reverse causal reasoning, the execution path is planned backwards from the target outcome, resulting in a more reasonable and efficient control sequence generated by traditional forward planning methods. Second, a three-layer progressive safety protection system is established. Negative space analysis technology proactively identifies prohibited dangerous action sets, providing preventive safety protection. Paradox verification scenario technology tests the system's tolerance under extreme conditions and determines control boundary parameters. Multiple interlocking verification conditions enable progressive safe unlocking of functions, ensuring that each operation is within the verified safety range. This "prevent-test-verify" active protection mechanism is more effective in preventing potential risks than traditional passive protection methods. Third, a complete remote control closed-loop control system is implemented. A parallel monitoring mechanism acquires multi-dimensional execution status data, enabling real-time command transmission and verification via wireless communication. The root causes of problems are identified through a layered, speed-dependent causal feedback chain, including a fast, strongly correlated core control layer, a slow, strongly correlated steady-state regulation layer, a fast, weakly correlated disturbance compensation layer, and a slow, weakly correlated trend prediction layer. Predictive correction technology adjusts control parameters in advance, reducing control deviations compared to traditional reactive control, enabling the robot to maintain stable control performance in complex environments.
[0009] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The accompanying drawings herein illustrate specific examples of the technical solutions described in the present invention, and together with the specific implementation methods constitute a part of the specification, and are used to explain the technical solutions, principles and effects of the present invention.
[0011] Unless otherwise specified or defined, the same reference numerals in different drawings represent the same or similar technical features, and the same or similar technical features may also be represented by different reference numerals.
[0012] Figure 1 The present invention is a flowchart of a method for remotely controlling a humanoid robot based on wireless communication.
[0013] Figure 2 This is a structural block diagram of a humanoid robot remote control device based on wireless communication of the present invention.
[0014] Figure 3 It is a structural schematic diagram of a computer device of the present invention. DETAILED DESCRIPTION
[0015] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0016] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0017] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0018] The technical solutions of the embodiments of this application are introduced below.
[0019] like Figure 1 As shown, an embodiment of the present invention provides a method for remotely controlling a humanoid robot based on wireless communication, comprising the following steps S110 to S170: Step S110 , receiving a remote control instruction from the operator, semantically parsing the remote control instruction to generate control intention data, establishing a two-way expected calibration channel based on the control intention data, and generating expected consistency parameters and deviation tolerance thresholds based on the two-way expected calibration channel.
[0020] Remote control command reception utilizes a multimodal input fusion architecture, integrating multiple control input methods, including voice commands, gesture recognition, brain-computer interface (BCI), and tactile feedback, for parallel processing. The voice command reception system utilizes a highly directional microphone array and a 16kHz sampling frequency for real-time audio capture, using noise suppression and echo cancellation to ensure command clarity. The gesture recognition system integrates an RGB-D depth camera and an IMU sensor to capture the operator's 3D hand motion trajectory in real time at a frame rate of 30fps. The BCI module utilizes a 64-channel EEG electrode array, sampling at a high frequency of 1000Hz to record EEG activity signals and extracts motor imagery features using a convolutional neural network. The tactile feedback device incorporates pressure sensors and vibration actuators to establish a bidirectional force interaction channel between the operator and the robot. All input signals are precisely time-aligned through hardware clock synchronization and software timestamp correction, and a distributed time protocol is employed to ensure the synchronous fusion of multimodal data. For example, in a home care scenario, when the operator issues a voice command "Help me lift up the elderly person sitting on the sofa", the multimodal fusion system simultaneously analyzes the urgency in the voice tone, the specific location of the gesture, and the intensity of the control intention reflected by the EEG signal to generate a complete remote control command.
[0021] Semantic parsing utilizes a hierarchical semantic understanding architecture. Through a multi-stage process involving lexical analysis, syntactic parsing, semantic reasoning, and intent recognition, it converts remote control commands into structured control intent data understandable by the robot. The natural language processing engine, based on a pre-trained language model using the Transformer architecture, undergoes domain-adaptive training on human-robot interaction data to build a specialized semantic knowledge base covering humanoid robot control scenarios. Speech signal processing utilizes an end-to-end speech recognition model, enabling a direct mapping from audio waveforms to semantic intent. Gesture semantic analysis builds a three-dimensional gesture dictionary and a temporal action pattern library, decomposing continuous hand movement sequences into discrete semantic primitives. EEG signal semantic decoding utilizes a hybrid architecture combining convolutional neural networks and long-short-term memory networks to extract stable representations of movement intent from the high-dimensional EEG feature space. The semantic fusion module employs an attention mechanism and graph neural network techniques to weight and reason about semantic fragments from different modalities, resolving conflicts and ambiguities between multimodal information. The context understanding mechanism maintains a dynamic dialogue state and task context, utilizing memory networks and knowledge graph reasoning to complete incomplete command expressions. For example, when the operator issues a vague command such as "robot come here", the semantic analysis system automatically infers the specific care needs and execution plan based on the current room layout, the elderly person's physical condition and historical care records, and generates complete structured control intention data.
[0022] A bidirectional expectation calibration channel is established based on the manipulation intent data generated by semantic analysis, realizing a framework for intelligently matching manipulation requirements with robot capabilities. Manipulation intent data is structured, containing key information such as task type, target object, action to be executed, and constraints. For example, the task of helping an elderly person stand up is represented as I={task:"assist_standing",object:"elderly",action:"lift",force_range:[50N,150N],speed:"slow",position:"sofa"}. The bidirectional calibration channel construction consists of two core components: expectation vector generation and capability vector generation. The expectation vector extracts quantitative requirements from the manipulation intent data, including dimensions such as expected assist force, expected speed, expected lifting angle, and expected completion time. The robot capability vector reflects actual technical indicators, including performance parameters such as maximum output force, maximum speed, joint range of motion, and fastest response time. The bidirectional calibration channel establishes a matching analysis mechanism, employs a multi-dimensional similarity assessment method, and designs a bidirectional message passing architecture with forward and backward propagation. The forward propagation path evaluates the degree to which the robot's capabilities meet the manipulation expectations, while the backward propagation path assesses the rationality of the manipulation expectations relative to the robot's capabilities. Through structured analysis based on manipulation intention data and the design of a two-way matching mechanism, a two-way expectation calibration channel was successfully established.
[0023] Based on the established bidirectional expectation calibration channel, an iterative negotiation mechanism generates expected consistency parameters and deviation tolerance thresholds. The expected consistency parameters are generated through an iterative negotiation process using bidirectional message passing and are dynamically updated based on the priority weights of various dimensions in the manipulation intention data. During the iterative process, the forward pass calculates the degree to which the robot's capabilities meet expectations, while the backward pass evaluates the rationality of the expectations. Through multiple iterative negotiations, the expected consistency parameters converge to the final value. For the "helping the elderly person" task, safety is prioritized, followed by comfort. After iterative negotiation through the bidirectional calibration channel, the expected consistency parameter converges to 0.83. The deviation tolerance threshold is also dynamically set based on the negotiation mechanism of the bidirectional calibration channel. First, basic constraints are extracted from the manipulation intention data. Then, bidirectional negotiation is conducted based on the actual robot's precision capabilities. When the robot's precision capabilities are high, the threshold is appropriately tightened; when capabilities are limited, the threshold range is relaxed. The expected execution result is generated using a parameter-driven mapping method: R_expect = EC × R_ideal + (1-EC) × R_safe. It is dynamically adjusted between the ideal and safe states based on the expected consistency parameters. Ultimately, a set of deviation tolerance thresholds is generated through a negotiation mechanism in a bidirectional calibration channel, including force deviation ±20N, velocity deviation ±0.05m / s, and angle deviation ±15°. This ensures that the execution process meets the intended requirements while maintaining a safety margin. The expected consistency parameters and deviation tolerance thresholds are ultimately generated through a bidirectional expectation calibration channel.
[0024] Step S120 , obtaining an expected execution result based on the expected consistency parameter, inferring an optimal control sequence based on the expected execution result, and performing negative space analysis on the optimal control sequence to identify a prohibited execution action set.
[0025] Specifically, the expected execution result is obtained based on the expected consistency parameter. The expected consistency parameter, EC, serves as a measure of the match between control expectations and robot capabilities and directly determines the criteria for setting the expected execution result. The expected execution result is derived using a parameter-driven mapping method: R_expect = EC × R_ideal + (1-EC) × R_safe, where R_expect is the expected execution result, R_ideal is the ideal state, R_safe is the safe and conservative result, and EC is the expected consistency parameter. When EC = 1.0, the expected execution result is equivalent to the ideal state; as EC decreases, the result automatically shifts toward the safe state. The deviation tolerance threshold T serves as a hard constraint to limit the acceptable range of the result, ensuring that |R_expect - R_actual| ≤ T. The expected execution result is represented as a state vector R = [p_target, θ_target, t_exec, f_max], which contains the target position p_target, target posture θ_target, execution time t_exec, and maximum force f_max. For example, in the transfer assistance task, when EC = 0.83, the expected performance results are: target position accuracy of ±5 cm, posture accuracy of ±10°, execution time of 8 seconds, and maximum assist force of 150 N. The ideal state requires higher accuracy (±2 cm, ±5°), while the safe state allows for a wider margin (±10 cm, ±20°). The expected performance results are achieved through linear mapping of the expected consistency parameters and threshold constraints.
[0026] In some embodiments, the reverse deduction of the optimal manipulation sequence based on the expected execution result includes: constructing a result state mapping table based on the expected execution result; reverse parsing the result state mapping table to obtain a causal association path; and reverse deducing the optimal manipulation sequence based on the causal association path.
[0027] Based on the expected execution result, a result-state mapping table is constructed, decomposing the target requirements in the expected execution result into a sequence of executable states. The expected execution result provides the task's endpoint constraints and process limits, requiring a complete mapping from the current state to the target state. The result-state mapping table is represented by a directed graph G = (V, E), where nodes V_i represent key states in the execution process and edges E_ij represent feasible transitions between states. Each state inherits constraints from the expected execution result, ensuring that S_i.f ≤ f_max and that the execution time satisfies Σt_ij ≤ t_exec. A state is defined as S_i = [p_i, θ_i, f_i, t_i], where the position p_i and the posture θ_i are gradually interpolated from the initial position to p_target and θ_target. The state discretization density is adjusted based on the accuracy requirements of the expected execution result, using dense sampling for high accuracy (EC > 0.8) and sparse sampling for low accuracy. The mapping weight w_ij reflects the priority of the state transition, with states closer to the expected execution result receiving higher weights. For example, the mapping table for the expected execution results of the lifting task includes key states such as initial approach, contact establishment, force application and ascent, position adjustment, stable maintenance, and force release. Each state satisfies the force constraint of f ≤ 150N and the total time constraint of 8 seconds. The result-state mapping table is generated by decomposing the expected execution results and propagating constraints.
[0028] The result state mapping table is reverse parsed, backtracking from the target state in the mapping table to identify causal paths. The result state mapping table G = (V, E) provides all possible state transition relationships. The reverse parsing requires finding the optimal path that satisfies the causal logic. The dynamic programming algorithm starts from the goal node V_goal and calculates the optimal cost V(s_i) = min{c(s_i,s_j) + V(s_j)} to reach each node. The transition cost c(s_i,s_j) directly uses the weights w_ij in the mapping table, taking into account the physical feasibility of state transitions. Causal constraints are extracted from the topology of the mapping table. For example, node V_contact (contact establishment) must precede node V_lift (lift). This dependency is represented in the mapping table as the existence of directed edges. During the path search, only paths that satisfy all constraints in the mapping table are retained. The backtracking algorithm, starting from V_goal, selects the predecessor node with the lowest cumulative cost until reaching the initial state V_start. Path verification checks whether the extracted path covers all required nodes in the mapping table. For example, the causal path extracted from the mapping table for the lifting task is: V_start → V_approach → V_contact → V_lift → V_adjust → V_stable → V_goal. Each transition corresponds to an edge in the mapping table. By reversely traversing the result state mapping table and verifying the constraints, the causal association path is obtained.
[0029] The optimal control sequence is generated by inversely factoring the causal path, converting the discrete state sequence in the path into continuous control commands. The causal path V_0→V_1→...→V_n provides the logical order and timing constraints for state transitions. Each state V_i contains target values such as position, attitude, and force. Trajectory generation interpolates between adjacent states V_i and V_i+1 to ensure smooth transitions. For each transition, a trajectory segment of the corresponding length is generated based on the time allocation t_i in the causal path. Critical transitions (such as contact establishment from V_approach to V_contact) are assigned more control points to ensure accurate execution. The interpolation method is selected based on the path characteristics, using cubic splines for position continuity and linear interpolation for force to avoid sudden changes. Joint space transformations map the Cartesian states S_i.p in the causal path to joint configurations q_i through inverse kinematics. Timing constraints ensure that the execution time of each trajectory segment strictly adheres to the plan in the causal path. Dynamic verification verifies that the generated trajectory satisfies the physical constraints implicit in the causal path. For example, based on the causal path of "approach → contact → ascent → stabilization," the generated control sequence maintains low speed during the contact phase (following the causal constraints) and constant contact force during the ascent phase (meeting the path requirements). By generating segmented trajectories based on the causal path and satisfying the constraints, an optimal control sequence is generated.
[0030] In some embodiments, the negative space analysis of the optimal control sequence to identify a set of prohibited actions includes: constructing a full action space matrix corresponding to the optimal control sequence; excluding the optimal control sequence from the full action space matrix to obtain a negative space area; analyzing the negative space area to generate a dangerous action marker; and obtaining a set of prohibited actions based on the dangerous action marker.
[0031] The full action space matrix corresponding to the optimal manipulation sequence is constructed. The complete set of reachable actions is determined based on the range of motion of the optimal manipulation sequence. The optimal manipulation sequence q_opt(t) defines the standard trajectory for task execution. The full action space must include all possible deviations from this trajectory. Matrix construction is centered on the optimal sequence and extended outward to the physical limits of the robot. At each time instant t_k along the optimal trajectory, spatial discretization is performed, and possible configurations are sampled within its neighborhood to form the matrix A[i,j,k]. The sampling density is adjusted based on the local characteristics of the optimal sequence, with additional sampling points added in areas of high trajectory curvature. Reachability assessment considers not only single-point reachability but also the ability to return after deviating from the optimal sequence. Boundary determination comprehensively considers joint limits, collision constraints, and task space constraints. The time dimension of the matrix corresponds to the discrete points of the optimal sequence to ensure temporal consistency. A sparse representation stores only the regions associated with the optimal sequence to improve storage efficiency. The full action space matrix is constructed through spatial expansion and constraint checking based on the optimal manipulation sequence.
[0032] The optimal maneuver sequence is excluded from the full action space matrix to obtain the negative space. The full action space matrix A contains all possible actions. The optimal maneuver sequence q_opt(t) occupies one of these trajectories, and the remaining negative space needs to be identified. This exclusion operation not only removes the optimal trajectory itself but also considers its safety envelope. This safety envelope is dynamically adjusted based on the velocity v_opt and acceleration a_opt of the optimal sequence, with a larger safety margin required at high speeds. Trajectory expansion uses a variable radius method, r(t) = r_base + k_v |v_opt(t)| + k_a |a_opt(t)|, to ensure that the safety distance varies with the motion state. The expanded trajectory area O_expand is subtracted from matrix A to obtain the negative space N = A - O_expand. Structural analysis of the negative space identifies topological relationships with the optimal sequence, including completely separated regions, adjacent regions, and enclosed regions. Connectivity analysis determines whether actions in the negative space can reach each other without traversing the optimal trajectory. The volume ratio V_N / V_A is calculated to quantify the extent to which the optimal sequence occupies the action space. For example, in the lifting task, the optimal sequence and its safety envelope occupy about 30% of the total action space, and the negative space is mainly distributed in the areas of rapid lateral movement and large swing.
[0033] Risk analysis is performed on the negative space region, assessing the dangerousness of the actions within it based on its characteristics. The negative space region N contains all actions that deviate from the optimal sequence, and the risk level of each action must be assessed individually. Risk assessment first analyzes the degree of deviation from the optimal sequence; greater deviations generally increase the risk. Collision risk is assessed by calculating the swept volume of the action in the negative space; a larger volume increases the probability of collision. Speed risk checks whether the speed of the action in the negative space exceeds the safety threshold v_safe. Torque risk analyzes whether the joint torque required to execute the action in the negative space is close to the limit. Stability risk is determined by checking whether the center of gravity trajectory exceeds the support range. Combined risk considers the cumulative effect of multiple risk factors and uses fuzzy logic to handle uncertainty. Risk levels are categorized as high, medium, and low based on the comprehensive score. The risk distribution map shows that high-risk areas in the negative space are primarily concentrated near the movement limits. For example, the analysis found that rapid swinging actions (located at the far end of the negative space) have a high collision risk, while actions that deviate slightly from the optimal trajectory have a lower risk. This multi-dimensional risk assessment of the negative space region generates a dangerous action marker.
[0034] Based on dangerous action tags, a set of prohibited actions is obtained, and actions marked as high-risk are added to the prohibited list. Dangerous action tags assign a risk level to each action in the negative space, and the prohibited set extracts those that exceed the safety threshold. The set construction rules are categorized by risk level: high-risk actions are directly classified as absolutely prohibited, medium-risk actions are conditionally prohibited, and low-risk actions are only prompted. Absolutely prohibited actions implement hard restrictions at the underlying controller level to ensure they are never executed. Conditionally prohibited actions determine whether they are permitted based on the specific scenario and safety measures. The parameterized description of prohibited actions includes joint configuration ranges, speed limits, acceleration limits, and more. The compact representation of the set eliminates redundancy, merging similar dangerous actions into a single category. A real-time update mechanism dynamically adjusts the prohibited boundaries based on task execution feedback. The index structure enables fast queries to determine whether a given action is in the prohibited set. For example, the final prohibited set for the elderly care scenario includes dangerous action categories such as rapid arm swings, sudden accelerations, and large-angle tilts. Through risk-based classification extraction and parameterized representation, a prohibited action set is established.
[0035] Step S130 , constructing a paradox verification scenario by combining the optimal control sequence and the prohibited action set, and performing a conflict test through the paradox verification scenario to generate control boundary parameters.
[0036] In some embodiments, the construction of a paradox verification scenario by combining the optimal manipulation sequence and the prohibited action set includes: performing intersection analysis on the optimal manipulation sequence and the prohibited action set to obtain an intersection analysis result; identifying potential conflict nodes based on the intersection analysis result; performing contradiction reinforcement processing on the potential conflict nodes to generate a conflict amplification factor; and using the conflict amplification factor to construct a paradox verification scenario.
[0037] Perform an intersection analysis on the optimal manipulation sequence and the set of prohibited actions to identify the overlapping and adjacent regions in the spatio-temporal domain. The intersection analysis uses the set operation I = Q_opt ∩ F_forbid, where Q_opt is the extended domain of the optimal sequence and F_forbid is the parameter space of the prohibited set. The direct intersection is usually empty (a well-designed controller will not plan the optimal path within the prohibited area), so neighborhood expansion analysis is required. The distance metric uses the weighted norm d = √(w_p||Δp||² + w_θ||Δθ||² + w_v||Δv||²), where w_p, w_θ, and w_v are the weight coefficients for position, angle, and velocity respectively. The neighborhood radius is determined based on the control accuracy and disturbance level, and typical values are 1.5 - 2 times the envelope of the optimal trajectory. The temporal correlation analysis examines the proximity of the optimal sequence to the prohibited actions at different times, generating a distance-time curve d(t). The intersection metric indicators include the minimum distance d_min, the average distance d_avg, the risk exposure time T_risk, etc. The boundary recognition algorithm extracts the common boundary between the optimal sequence and the prohibited area, and these positions are the most likely to generate conflicts. The intersection visualization maps the results to the configuration space and shows the distribution of the conflict intensity using a heat map. For example, analysis finds that at t = 3.5 s for the lifting action, the elbow joint angle is only 8° away from the boundary of the prohibited area, constituting a potential risk point. Through multi-dimensional set operations and distance analysis, detailed intersection analysis results are obtained.
[0038] Identify potential conflict nodes based on the intersection analysis results, and locate the key positions and times where contradictions are most likely to occur. The intersection analysis results provide the distance distribution d(t) and boundary information between the optimal sequence and the prohibited area, and the conflict node identification conducts risk assessment based on this. The node is defined as N_c={(q_i,t_i)|d(q_i,F_forbid)<d_threshold}, where d_threshold is the risk distance threshold. The risk score comprehensively considers the static proximity and dynamic trend: R_node = α / d + β|dd / dt|, and the risk increases sharply when the trajectory approaches the prohibited area rapidly. The key node screening retains several points with the highest risk scores, usually choosing the local maximum points. The influence range of the node is determined through sensitivity analysis, calculating the impact of the perturbation at this point on the overall trajectory. The time clustering merges adjacent conflict nodes to avoid over-segmentation. The node type classification includes: approaching type (small static distance), crossing type (dynamic trend pointing to the prohibited area), and oscillating type (repeatedly near the boundary). The causal chain analysis traces the formation reasons of each conflict node, whether it is caused by task requirements or constraint limitations. For example, 3 main conflict nodes are identified: the speed in the rapid rising section approaches the upper limit, the angle during turning approaches the singular configuration, and the force control accuracy requirement during the contact stage is extremely high. Through risk scoring and multi-dimensional analysis, potential conflict nodes are accurately identified.
[0039] Potential conflict nodes are subjected to conflict intensification, artificially amplifying conflict characteristics to fully expose weaknesses in the control scheme. Conflict nodes N_c include the locations most prone to problems. Conflict intensification transforms potential conflicts into explicit conflicts by adjusting parameters. Intensification strategies include constraint tightening and demand amplification. Constraint tightening restricts the bounds from shrinking inward by Δd, while demand amplification increases the expected performance by Δp. The conflict amplification factor λ_i is designed for each node individually, reflecting its conflict sensitivity. The calculation formula is λ_i = 1 + k_risk × R_node_i, where k_risk is the risk factor. High-risk nodes receive a larger amplification factor. Parameter perturbations adopt a normal distribution N(μ,σ²), with mean μ = λ_i × p_nominal and standard deviation σ reflecting uncertainty. Temporal intensification is achieved by compressing execution time, requiring the same action to be completed in a shorter time. Spatial intensification reduces the available space to simulate operations in confined environments. Combined intensification adjusts multiple parameters simultaneously to generate complex conflicts. Intensified bounds prevent physically impossible scenarios, maintaining the practical relevance of the test. For example, for speed-critical nodes, the maximum speed limit was reduced by 20% (λ = 0.8) while maintaining the original execution time, creating a speed-time conflict. Through systematic parameter adjustment and conflict amplification, an effective conflict amplification factor was generated.
[0040] Conflict amplification factors are used to construct complete paradox verification scenarios, creating an executable test environment. The conflict amplification factor λ_i provides the degree of conflict intensification at each node. Scenario construction requires integrating these local conflicts into coherent test cases. Scenario parameters are obtained by multiplying the original parameters by the amplification factor: p_scenario = p_original × λ. Constraint updates take into account the propagation effect of conflict; the strengthening of one node may affect adjacent nodes. The spatiotemporal description of the scenario includes: initial conditions (the strengthened starting state), target conditions (the improved performance requirements), constraint sets (the tightened constraints), and a disturbance model (simulating external perturbations). Executability checks ensure that the constructed scenarios have theoretical solutions; scenarios with no solutions are not worth testing. Scenario difficulty is graded based on the degree of conflict: mild paradox (a difficult but feasible solution exists), moderate paradox (a solution requires significant compromise), and severe paradox (a near-unsolvable extreme). Dynamic scenarios introduce a time-varying conflict factor λ(t) to simulate changes in conflict intensity. The scenario library includes multiple typical paradoxes, covering different types of conflict. For example, in a "rapid transfer in a confined space" scenario, the workspace was reduced by 30%, while the execution time requirement was shortened by 25%, all while maintaining the original safety constraints. This created a triple contradiction between space, time, and safety. Through the systematic integration of conflict factors and scenario parameterization, a paradox verification scenario was successfully constructed.
[0041] In some embodiments, the generating of control boundary parameters by performing a conflict test through the paradox verification scenario includes: simulating the execution of contradictory instructions in the paradox verification scenario; analyzing the execution process of the contradictory instructions to obtain conflict response data; performing extreme value analysis based on the conflict response data to determine a tolerance threshold; and obtaining control boundary parameters according to the tolerance threshold.
[0042] Paradox verification scenarios simulate the execution of contradictory instructions to test the robot's actual performance under conflicting conditions. Paradox verification scenarios provide strengthened constraints and requirements. The generation of contradictory instructions requires finding a feasible solution under these conflicting conditions. Instruction generation utilizes a multi-objective optimization approach, with the objective function consisting of a performance metric and a constraint violation penalty: J = w_perf × P + w_viol × V, where P is the performance score and V is the constraint violation penalty. The optimization algorithm employs sequential quadratic programming (SQP) to handle nonlinear constraints. When a solution that fully satisfies all conditions cannot be found, the algorithm returns a suboptimal solution with the smallest constraint violation. Execution monitoring records time series of key variables, such as position error e_p(t), force overrun f_over(t), and velocity violation v_viol(t). Conflict trigger detection identifies the precise time and location of constraint violations. Multiple executions employ Monte Carlo methods, introducing random perturbations to test robustness. Execution strategies include forced execution (ignoring some constraints), adaptive execution (dynamically adjusting parameters), and segmented execution (using different strategies at different stages). For example, in a confined space scenario, simulation execution showed that the collision warning was first triggered at t=2.3s. At t=3.7s, the speed constraint was violated, forcing the robot to slow down and causing a mission timeout. By executing conflicting instructions and monitoring the entire process, we captured data on the robot's behavior under conflicting conditions.
[0043] The execution process of conflicting instructions is analyzed to extract the robot's response characteristics under conflict conditions. The time series data generated during the execution process contains rich conflict response information, and key patterns can be identified through signal processing and feature extraction. Conflict identification algorithms detect constraint violations and calculate the frequency, magnitude, and duration of violations. Response delay analysis measures the time difference Δt_response from the occurrence of a conflict to the controller's adjustment, reflecting the controller's response speed. Oscillation detection uses spectrum analysis to identify high-frequency components in the control signal. Excessive oscillation indicates that the control is approaching instability. The performance degradation curve P(t) shows the downward trend in execution performance as the conflict intensifies. Constraint prioritization observes which constraints are relaxed first to infer the inherent priority settings of the control strategy. Failure mode classification categorizes observed problems into: graceful degradation (performance degradation but controllable), oscillation divergence (control instability), and hard failure (complete failure). Stress concentration analysis identifies the weakest link in the control scheme, often manifested as a variable reaching its limit first. Resilience assesses the robot's ability to recover from a conflict. For example, analysis revealed that position control prioritized accuracy during conflicts, automatically reducing speed; force control exhibited 0.5Hz oscillations, indicating approaching the stability boundary; and control strategies tended to sacrifice time performance to meet safety constraints. Through multi-dimensional signal analysis and pattern recognition, detailed conflict response data was obtained.
[0044] Extreme value analysis is performed based on conflict response data to determine the maximum conflict intensity that the control scheme can withstand. Conflict response data reveals the robot's performance under different conflict levels. Extreme value analysis aims to identify acceptable performance boundaries. The performance metric is defined as a comprehensive score of key parameters: S = Σw_i × (1 - |p_i - p_target| / p_range). When S falls below a threshold, performance is considered unacceptable. Extreme value search uses a binary search method to find the critical point λ_critical within the conflict intensity range [λ_min, λ_max]. The tolerance curve plots the extreme values under different conflict types, forming a multidimensional tolerance boundary. Statistical analysis considers the distribution characteristics of the response data and determines a conservative tolerance threshold using a 95% confidence interval. The failure probability model P_fail(λ) = 1 / (1 + exp(-k(λ - λ_critical))) describes the relationship between conflict intensity and failure probability. Safety margins are set to ensure that control maintains a certain buffer even near extreme values. Sensitivity analysis identifies the parameters that have the greatest impact on extreme values, which require key monitoring. Time-varying extreme values account for actuator fatigue, which can degrade tolerance after long-term operation. For example, analysis determined that the tolerance threshold for position conflict is a ±12% trajectory deviation, exceeding which the position error increases dramatically; the limit for velocity conflict is 75% of the nominal value; any lower than this will result in mission failure; and the oscillation amplitude of force control should not exceed 20% of the expected value. Through systematic extreme value search and statistical analysis, the tolerance thresholds for each dimension were determined.
[0045] Based on the tolerance threshold, control boundary parameters are derived to define the safe and feasible operating range. The tolerance threshold provides the robot's ultimate capabilities under various collision scenarios. Based on this, reasonable safety margins are set to derive control boundary parameters. Boundary parameters are defined as B = {b_pos, b_vel, b_force, b_time}, corresponding to position, velocity, force, and time constraints, respectively. The safety factor k_safe is chosen to balance conservatism and usability, with a typical value of 0.7-0.8, meaning the boundary is set at 70-80% of the tolerance threshold. The coupling relationship between parameters is determined through correlation analysis, and some parameters require coordinated adjustment. Boundaries are categorized as hard and soft: hard boundaries must not be crossed, while soft boundaries allow for temporary crossings. Dynamic boundaries adjust based on the current state, automatically tightening other boundaries when a limit is detected. Boundary visualization draws a safety envelope in the operating space, visually displaying the feasible area. Validation testing ensures that the set boundaries actually avoid issues observed in collision scenarios. The structured representation of boundary parameters supports fast query and real-time judgment. For example, the final control boundary parameters are: position boundary b_pos = ±8 cm (based on a 12% tolerance threshold), velocity boundary b_vel = [0.1, 0.8] m / s, force boundary b_force = [10, 120] N, and time margin b_time = 1.2 × t_nominal. Through conservative design based on the tolerance threshold and multi-dimensional parameter definition, the complete control boundary parameters were finally generated.
[0046] Step S140 : establishing an interlock verification condition based on the control boundary parameter and the deviation tolerance threshold, and generating a safe unlocking sequence based on the interlock verification condition.
[0047] In some embodiments, establishing interlocking verification conditions based on the control boundary parameters and the deviation tolerance threshold includes: performing a correlation analysis on the control boundary parameters and the deviation tolerance threshold to obtain a correlation analysis result; constructing a multidimensional verification matrix based on the correlation analysis result; performing constraint condition screening on the multidimensional verification matrix to obtain key constraint factors; and establishing interlocking verification conditions based on the key constraint factors.
[0048] Perform a correlation analysis between the manipulation boundary parameters and the deviation tolerance threshold to reveal the internal relationship between the two parameter systems. Although the manipulation boundary parameter B and the deviation tolerance threshold T have different sources, they influence each other during the control process. The correlation calculation uses the mutual information method I(B,T)=ΣΣp(b,t)log(p(b,t) / (p(b)p(t))) to quantify the information dependence between parameters. The linear correlation analysis calculates the Pearson coefficient r_ij=cov(B_i,T_j) / (σ_Bi×σ_Tj) to identify direct linear relationships. The non-linear correlation is analyzed after mapping to a high-dimensional space through a kernel function to capture complex dependence patterns. The time-delay correlation considers the lag effect of parameter influence, where changes in some boundary parameters affect the effectiveness of the deviation threshold after a delay τ. The cross-influence analysis identifies the mutual constraints between parameters, such as tightening the position boundary increases the sensitivity to angular deviation. The correlation strength grading classifies parameter pairs into strong correlation (r>0.7), medium correlation (0.3<r≤0.7), and weak correlation (r≤0.3). The dynamic correlation tracks the changes in parameter relationships over the execution process. The correlation map visually displays the parameter network, with the node size representing importance and the edge thickness representing the correlation strength. For example, analysis reveals that the position boundary b_pos and the position deviation threshold T_pos are strongly positively correlated (r=0.85), and the velocity boundary b_vel and the time deviation threshold T_time are moderately negatively correlated (r=-0.52). Through the comprehensive application of various correlation analysis methods, detailed correlation analysis results are obtained.
[0049] A multidimensional validation matrix is constructed based on the results of correlation analysis, forming a structured validation framework. The correlation analysis results provide the relationship strength r_ij and influence patterns between parameters. The validation matrix organizes these relationships into an actionable form. The matrix dimensions include boundary parameter dimensions (4 dimensions), deviation threshold dimensions (4 dimensions), and validation type dimensions (three categories: range check, trend check, and combination check). The matrix element V[i,j,k] defines the validation rule for the i-th boundary parameter and the j-th deviation threshold under the k-th validation category. Rule generation is determined by correlation strength. Strongly correlated parameter pairs are subject to strict joint validation, while weakly correlated parameters are subject to independent validation. The validation function design considers the type of correlation: positive correlations are subject to unidirectional constraints, while negative correlations are subject to inverse compensation. Weight assignment is based on correlation strength: w_ij = |r_ij| / Σ|r_ij|, ensuring that important correlations receive more attention. Sparse processing removes validation items with correlations below the threshold (|r| < 0.1) to improve computational efficiency. The block structure groups highly correlated parameters into validation blocks, with joint validation within the blocks and independent processing between blocks. The time-varying nature is reflected by introducing the time dimension t. V[i, j, k, t] allows verification rules to evolve over time. For example, in the constructed verification matrix, the position-force coupling block contains four strongly correlated verification rules, the position-velocity coupling block contains three strongly correlated verification rules, and the velocity-time block contains three moderately correlated rules. Through the structured organization of correlations and rule generation, a multidimensional verification matrix is constructed.
[0050] Constraint screening is performed on the multidimensional verification matrix to identify the key factors with the greatest impact on safety. The multidimensional verification matrix V contains a large number of verification rules, and the most critical constraints need to be extracted through screening. Importance assessment uses sensitivity analysis to calculate the change in safety indicators (ΔS) after removing a verification rule. Redundancy detection identifies duplicate verification rules and retains the most stringent one. Constraint propagation analyzes the logical implications between verification rules, eliminating those implied by other rules. Principal component analysis reduces the high-dimensional verification space; the top three to four principal components typically contain more than 80% of the verification information. The critical path identifies the necessary verification nodes from the initial state to the critical state; these nodes correspond to critical constraints. Failure mode and effects analysis (FMEA) assesses the severity of the consequences of each constraint failure. Historical data is statistically analyzed to determine the triggering frequency and interception effectiveness of each constraint. The screening criteria comprehensively consider importance, uniqueness, and effectiveness, with a score of S = α × Importance + β × Uniqueness + γ × Effectiveness. A threshold is set to retain the top 70% of constraints as critical constraints. For example, the screening results show that position-velocity coupling constraints, force-time gradient constraints, and angle-position coordination constraints are the three most critical constraint factors. Through multi-dimensional analysis and scientific screening methods, the key constraint factors were obtained.
[0051] Interlock verification conditions are established based on key constraints to form an executable safety assurance mechanism. Key constraints represent the most important safety requirements, and interlock verification conditions must translate these factors into specific judgment logic. Conditions are expressed using Boolean logic, with each key constraint corresponding to one or more logical expressions. Basic conditions include range checks (x_min ≤ x ≤ x_max), rate-of-change limits (|dx / dt| ≤ v_max), and combination constraints (f(x, y) ≤ threshold). Compound conditions are connected using logical operators: C_total = C1 AND (C2 OR C3) AND NOT C4. Temporal logic handles the order of conditions, using temporal operators to represent time relationships such as "before," "after," and "during." Trigger mechanisms define the timing of condition checks, including periodic, event, and conditional triggers. Priority settings ensure that critical conditions are checked first, while less critical conditions are processed later. Interlock depth controls the stringency of conditions and can be dynamically adjusted based on risk level. Failure handling specifies the response actions when conditions are not met, such as limiting functionality, degrading operation, or performing an emergency stop. The monitorable design of conditions ensures that all conditions can be verified through sensors or calculations. For example, the core interlock condition is: IF (|p_actual - p_target| > 0.8 × T_pos) AND (v_actual > 0.5 × b_vel) THEN lock_high_speed_motion. Through the design of logical expressions and the definition of execution mechanisms, a complete set of interlock verification conditions is established.
[0052] Based on the established interlock verification conditions, a safe unlock sequence is generated, and a progressive permission release mechanism is designed. The interlock verification conditions define various lock states and unlock conditions. A safe unlock sequence requires determining the appropriate unlocking order and timing. The unlock sequence is represented using a state machine model, with each state corresponding to a set of active lock conditions, and state transitions representing unlocking operations. Sequence generation adheres to the principle of least privilege, unlocking functions only when necessary. Unlocking priorities are determined based on mission requirements and safety levels, with basic motion functions unlocked first and high-risk operations released last. Timing constraints ensure that unlocking operations are not performed too quickly, with a minimum interval Δt_unlock set between successive unlocks. Condition monitoring continuously checks the execution effect after unlocking, and immediately reverts to a safe state if an anomaly occurs. An unlock token mechanism prevents unauthorized unlocking operations, with each unlocking step verifying a specific combination of conditions. Progressive testing performs small movements after each unlock to verify the controller's response. Sequence integrity checks ensure that all necessary functions are ultimately unlocked while maintaining safety constraints. For example, a typical unlocking sequence: Initially lock all degrees of freedom → Unlock position monitoring after verifying sensor function → Unlock and move slowly after confirming the environment is safe → Unlock force control after establishing stable contact → Unlock and move at full speed after all conditions are met. A state machine design and a progressive unlocking strategy generate a safe unlocking sequence.
[0053] Step S150 , performing parallel monitoring on the optimal control sequence based on the security unlocking sequence to obtain execution verification data, determining a target execution branch according to the execution verification data, and generating a unified instruction stream based on the target execution branch.
[0054] In some embodiments, the parallel monitoring of the optimal manipulation sequence based on the safe unlocking sequence to obtain execution verification data includes: decomposing the safe unlocking sequence into multiple monitoring checkpoints; performing real-time state sampling of the optimal manipulation sequence at the monitoring checkpoints to obtain sampling state data; analyzing the sampling state data to generate a state change trend; and generating execution verification data based on the state change trend.
[0055] The safe unlock sequence is broken down into multiple monitoring checkpoints, with state observation locations set at key nodes in the unlock process. The safe unlock sequence encompasses a state transition sequence S_unlock = {s_0, s_1, ..., s_n} from full lock to full functional release. Each state transition point represents a change in control authority. Checkpoints are set using a risk-driven strategy, with dense placement before and after high-risk unlocking operations and sparser placement during low-risk phases. Checkpoint types include mandatory checkpoints (must be passed before each unlock), conditional checkpoints (triggered under specific conditions), and periodic checkpoints (at fixed intervals). The timestamp T_check = {t_0, t_1, ..., t_m} records the expected trigger time of each checkpoint, where t_i+1-t_i ≥ Δt_min to ensure that check intervals are not too frequent. Check content is defined as C_i = {params_i, conditions_i, actions_i}, including the set of parameters to be monitored, the judgment conditions, and the response actions. Priority assignment ensures that critical checkpoints are executed first, with P_check∈{HIGH,MEDIUM,LOW}. The spatial distribution of checkpoints takes into account the trajectory characteristics of the optimal control sequence, increasing checkpoint density in locations with significant trajectory curvature and rapid speed changes. A dynamic adjustment mechanism allows for the addition and removal of checkpoints based on real-time status. For example, in the unlocking sequence for the assisted standing task, eight checkpoints are set: initial state confirmation, sensor unlock verification, slow movement test, contact establishment verification, force-controlled unlock verification, coordinated motion test, full-speed unlock verification, and task completion verification. Through scientific checkpoint placement and parameter design, the monitoring decomposition of the safe unlocking sequence is achieved.
[0056] Real-time state sampling of the optimal control sequence is performed at monitoring checkpoints to capture the instantaneous characteristics of the execution process. The optimal control sequence q_opt(t) defines the ideal execution trajectory. At each checkpoint t_i, the actual state q_actual(t_i) is collected for comparison. Sampled data includes multi-dimensional information such as joint configuration vector q, Cartesian position p, velocity v, acceleration a, and contact force f. A sampling synchronization mechanism ensures the temporal consistency of all sensor data, combining hardware triggering and timestamp correction. High-frequency sampling expands the time window Δt_window before and after the checkpoint, with a sampling rate of f_s = 1kHz, to capture detailed characteristics of the dynamic process. Data preprocessing includes outlier removal, noise filtering, and unit normalization to ensure data quality. State integrity checks verify the validity of the sampled data, and resampling is triggered if the missing data rate exceeds 5%. The deviation calculation e_i = q_actual(t_i) - q_opt(t_i) quantifies the difference between the actual and ideal values. Multi-sensor fusion uses weighted least squares to integrate information from multiple sources, including encoders, IMUs, and force sensors. A caching mechanism stores data 100ms before and after a checkpoint, supporting post-analysis. Sampling trigger conditions include time trigger (reaching t_i), event trigger (unlocking completed), and state trigger (deviation exceeded). For example, at the contact establishment checkpoint, sampled data shows: position p_actual = [0.52, 0.18, 0.35]m, a deviation of 2.3cm from the expected position; contact force f_actual = 45N, within the expected range of [40, 60]N; and joint velocities are all below 0.3rad / s, meeting safety requirements. Sampled state data is obtained through high-precision real-time sampling and multi-dimensional state capture.
[0057] In-depth analysis of sampled state data identifies state evolution patterns during execution. Time series analysis is used to extract trends from the discrete state points {q(t_0), q(t_1), ..., q(t_m)} obtained through sampling. Trend fitting uses a least-squares polynomial: q_trend(t) = a_0 + a_1t + a_2t², with the coefficient a derived through regression of the most recent N sampling points. Rate-of-change calculations (dq / dt) and d²q / dt²) reflect the dynamic characteristics of the state. The first-order derivative indicates velocity trends, and the second-order derivative indicates acceleration trends. Spectral analysis uses FFT transforms to identify periodic oscillations. A dominant frequency f_dominant > 5Hz indicates possible control instability. Phase space reconstruction embeds time series data into a high-dimensional space, revealing attractor structures and chaotic characteristics. Trend prediction uses the ARIMA model to predict the state k steps into the future based on historical data: q(t+k) = φ_1q(t) + φ_2q(t-1) + ... + ε(t). Abnormal pattern recognition includes mutation detection (|dq / dt| > threshold), drift detection (continuously increasing cumulative deviation), and oscillation detection (frequent alternation between positive and negative). Trend classification categorizes identified patterns into: convergence (gradually decreasing deviation), divergence (continuously increasing deviation), stability (constant deviation), and oscillation (periodic variation in deviation). Statistical feature extraction, including mean, variance, skewness, and kurtosis, comprehensively describes the state distribution characteristics. For example, analysis revealed that position deviation exhibited a convergence trend with a convergence time constant of τ = 2.5s; velocity exhibited small oscillations of 0.8Hz with an amplitude of 0.05 rad / s; and force control was stable with a variance of σ² = 2.3N². Through the combined application of multiple analytical methods, detailed state change trends were generated.
[0058] Execution verification data is generated based on state change trends. State change trends reveal the dynamic characteristics of the execution process, which need to be converted into verification data that can be used for decision-making. The verification data structure V_data = {V_safety, V_performance, V_stability, V_convergence} corresponds to safety, performance, stability, and convergence verification, respectively. Safety verification V_safety checks whether all states are within the safety boundary: s_safe = min(1, d_boundary / d_actual), where d_boundary is the distance to the danger zone. Performance verification V_performance evaluates execution quality: p_score = 1-Σw_i|e_i| / e_max, which integrates multi-dimensional deviations such as position, velocity, and force. Stability verification V_stability is based on the Lyapunov theory: if V(x) = x^TPx is decreasing and dV / dt < 0, the system is locally stable. Convergence verification V_convergence determines whether the system is moving toward the target: convergence time is estimated by fitting the deviation curve e(t) = e_0exp(-t / τ). Confidence calculation is based on the number of samples and data consistency: conf = 1 - σ_data / μ_data × √(n_min / n_actual). Timeliness marking ensures the freshness of validation data; data exceeding the validity period T_valid requires updating. Validation reports include numerical results, trend charts, risk warnings, and recommended actions. Anomaly annotation specifically identifies indicators that fall outside the normal range and analyzes the causes. For example, the generated validation data shows: safety 0.92 (good), performance 0.78 (improvement needed), stability 0.95 (excellent), convergence 0.81 (expected convergence within 3.2 seconds), and overall confidence 0.87. Through systematic data organization and multi-dimensional verification, complete execution validation data is generated.
[0059] The target execution branch is determined based on execution verification data, selecting the control strategy that best suits the current state. Execution verification data V_data provides a comprehensive assessment of the current execution state. Target branch selection requires finding the optimal match within a predefined set of strategies. The execution branch definition B = {B_nominal, B_conservative, B_aggressive, B_recovery} corresponds to the four strategies: standard execution, conservative execution, aggressive execution, and recovery execution, respectively. The branch selection function f_select: V_data → B is based on a comprehensive score from the verification data. The decision rule employs fuzzy logic: IF safety IS high AND performance IS low THEN B_aggressive, allowing performance to be optimized while preserving safety. Priority judgment ensures safety always takes precedence: when V_safety < 0.7, B_conservative is forced to be selected. The state machine model defines transition conditions between branches to prevent instability caused by frequent switching. Each branch includes a specific parameter adjustment strategy: B_conservative reduces speed by 50% and increases redundant checks, while B_aggressive increases speed by 30% and reduces intermediate verification. Historical performance learning records the execution results of each branch to dynamically adjust the selection bias. Weigh multiple objectives to find a balance between security, performance, and energy consumption. Through scientific decision-making mechanisms and strategy matching, we determine the target execution branch.
[0060] A unified instruction stream is generated based on the target execution branch, converting the branch strategy into an executable control command sequence. The target execution branch B_selected determines the control strategy, which must be combined with the optimal control sequence to generate specific instructions. The instruction stream generation uses a policy superposition method: I_final = I_optimal + ΔI_branch, where I_optimal is the original optimal instruction and ΔI_branch is the correction to the branch strategy. Parameter mapping converts the abstract branch strategy into concrete values: velocity scaling factor k_v, position offset δ_p, force adjustment Δ_f, etc. Timing alignment ensures temporal continuity of the corrected instruction stream, filling any gaps through interpolation. The instruction format is unified into {timestamp, joint_angles, velocities, torques, flags} to facilitate parsing by the underlying controller. Smoothing prevents sudden changes in instructions, using an S-shaped transition curve: x(t) = x_0 + (x_1 - x_0) × (3t² - 2t³). Integrity checking verifies that the instruction stream covers the entire execution cycle, without omissions or duplications. Priority marking distinguishes critical instructions from auxiliary instructions, ensuring that important instructions are executed first. The buffer design allows for a certain degree of execution flexibility, allowing non-critical instructions to be delayed or skipped. Real-time performance is guaranteed through timestamps and delay compensation mechanisms. Through policy mapping and instruction organization, a unified instruction stream is ultimately generated.
[0061] Step S160: The unified instruction stream is transmitted to the robot side via wireless communication. The robot side performs instruction verification based on the safety unlocking sequence to generate a verification pass signal. Based on the verification pass signal, the robot is driven to execute and obtain actual execution status data.
[0062] Specifically, a unified command stream is transmitted to the robot via wireless communication to ensure reliable delivery of control commands. The unified command stream (I_unified) contains structured information such as timestamps, control parameters, and priority tags, which require encoding and encapsulation for wireless transmission. The communication protocol uses the Real-Time Control Protocol (RCP). UDP is used at the transport layer to ensure low latency, and sequence numbers and checksums are added at the application layer to ensure reliability. The data frame format is [Header|Sequence|Timestamp|Priority|Payload|CRC], ensuring data integrity and timing accuracy. Transmission optimization uses differential encoding to reduce data volume. For consecutive similar commands, only the modified portion is transmitted. A priority queue mechanism ensures that critical commands, such as emergency stop, are transmitted first. The wireless link uses the 5G millimeter wave band, providing <5ms latency and >99.99% reliability. Channel adaptation dynamically adjusts the modulation and coding scheme based on RSSI and SNR. Redundant transmission uses spatial diversity to improve reliability for high-priority commands. Encryption uses AES-128 to ensure data security, with regular key updates. For example, a unified stream containing 500 instructions is transmitted: the original 50KB of data is differentially encoded and compressed to 15KB, which is then divided into 15 1KB data frames. Critical instructions are marked with the highest priority and sent immediately, while common instructions are transmitted sequentially. Transmission is completed over a 5G link in 1.5ms. After CRC verification at the receiving end, the data is sent to the verification module. Through optimized protocol design and transmission strategies, efficient and reliable transmission of unified instruction streams is achieved.
[0063] On the robot side, command verification is performed based on the safety unlock sequence to ensure that received commands comply with the current safety state. The safety unlock sequence S_unlock defines a progressive release of functionality, including states such as {s_lock, s_sense, s_slow, s_contact, s_force, and s_full}. Each state corresponds to a specific set of permissions. The verification process first analyzes the command type, parameters, and execution requirements. The permission mapping table M_permission establishes the mapping between commands and permissions. For example, high-speed motion requires the "full_speed" permission, and force control requires the "force_control" permission. The current permission P_current is derived from the execution progress of the unlock sequence. The verification algorithm checks each command: ifRequired_Permission(i) ⊆ P_currentthenValidelseInvalid. Timing verification ensures that unlocked functions are not executed prematurely. Parameter range checks verify that the parameters are within the permitted range for the current level, such as speed v ≤ v_max(s_current). Integrity verification checks dependencies between commands to prevent permission restrictions from disrupting execution continuity. Conflict detection identifies command sequences that are individually safe but combined and dangerous. Verification results are categorized as PASS (direct execution), MODIFY (execution after adjustment), WAIT (waiting for unlocking), and REJECT. Intelligent correction automatically adjusts out-of-limit parameters to a safe range. For example, if a 150N force control command is received in the s_contact state and the current limit is 100N, the verification module automatically adjusts it to 100N and generates a warning, ensuring safety while maximizing user intent. Through multi-level permission verification and intelligent processing, a verification pass signal is generated.
[0064] Based on the verification pass signal, the robot is driven to execute and obtain the actual execution status data. The verification pass signal V_pass contains the executable instruction queue and monitoring parameters, which are translated into the underlying control command by the actuator interface layer. The real-time controller runs at a frequency of 1kHz and completes the following in each cycle: sensor reading → state estimation → control calculation → instruction output → data recording. Multi-axis coordination is achieved through the EtherCAT bus at the μs level. Execution monitoring collects multi-dimensional status data: joint position q is obtained through a high-resolution encoder (accuracy 0.01°), speed q -The contact force F is calculated using differential filtering. A 6-axis force sensor (with an accuracy of 0.1N) is used for measurement, and temperature T is monitored to prevent overheating. Status data is formatted as [timestamp|joint_states|force|temperature|status], using efficient binary encoding. Anomaly detection rules include position overrun, velocity overrun, force overrun, and overtemperature, triggering graded protection responses. Performance metrics such as trajectory tracking error, steady-state accuracy, and response time are calculated in real time. A circular buffer is used for data caching, capable of storing 10 seconds of historical data. For example, during a grasping action, 1000 sets of position data are collected per second during the approach phase, displaying a smooth trajectory. A contact force of 28N is detected at the moment of contact, triggering a mode switch. During the grasping phase, a constant force of 30±2N is maintained, with a position accuracy of ±0.5mm. The overall tracking error is 1.8mm RMS, with a maximum error of 3.2mm, and no overrun alarms. Through comprehensive status monitoring and high-frequency data acquisition, detailed data on the actual execution status is ultimately obtained.
[0065] Step S170 , performing convergence analysis based on the actual execution state data and the deviation tolerance threshold to generate a convergence monitoring result, obtaining a causal feedback chain based on the convergence monitoring result, and using the causal feedback chain to predictively correct subsequent control behaviors to complete remote control.
[0066] Specifically, convergence analysis is performed based on the actual execution status data and the deviation tolerance threshold to evaluate the stability of the execution process and the target approximation characteristics. The actual execution status data D_actual includes multi-dimensional time-series information such as position, velocity, and force. The deviation tolerance threshold T defines the allowable deviation range for each dimension. The convergence analysis first calculates the real-time deviation e(t) = x_actual(t) - x_target, where x represents each state variable. The deviation evolution trend is obtained by exponential fitting e(t) = e_0×exp(-t / τ), and the convergence time constant τ reflects the speed of convergence. The convergence criterion is set as |e(t)| < T for a duration exceeding t_stable, indicating that the system enters a stable convergence state. The multi-dimensional convergence evaluation analyzes the convergence characteristics of each degree of freedom respectively to identify the bottleneck dimension with the slowest convergence. The convergence rate is calculated as r = -de / dt / e to quantify the instantaneous convergence speed. Oscillation detection determines whether there is a limit cycle oscillation by the number of zero crossings and the energy distribution. The convergence domain estimation uses the Lyapunov method, V(e) = e^TPe, and if dV / dt < 0, it is locally convergent. Statistical characteristics analyze the consistency of the convergence process and calculate the distribution of convergence parameters for different execution cycles. Abnormal pattern recognition includes divergence (τ < 0), oscillation (periodic), stagnation (r ≈ 0), etc. For example, the analysis shows that the position convergence time constant τ_p = 2.5s, the force convergence τ_f = 1.8s, and the overall converges to the threshold range within 4 seconds; however, a small oscillation of 0.5Hz is detected in the velocity dimension and needs attention. Through multi-dimensional convergence analysis, detailed convergence monitoring results are generated.
[0067] In some embodiments, obtaining the causal feedback chain according to the convergence monitoring results includes: performing hierarchical causal relationship identification on the convergence monitoring results to obtain strong correlation factors and weak correlation factors; performing temporal and spatial characteristic analysis based on the strong correlation factors and weak correlation factors, subdividing the strong correlation factors into fast strong correlation and slow strong correlation, and subdividing the weak correlation factors into fast weak correlation and slow weak correlation to obtain four types of spatio-temporal correlation factors; constructing a multi-level feedback network with hierarchical speed based on the four types of spatio-temporal correlation factors, where the fast strong correlation forms the core control layer, the slow strong correlation forms the steady-state regulation layer, the fast weak correlation forms the perturbation compensation layer, and the slow weak correlation forms the trend prediction layer; performing path search with priorities based on the multi-level feedback network to generate a causal feedback chain with hierarchical spatio-temporal characteristics.
[0068] Convergence monitoring results are used to identify hierarchical causal relationships and extract causal correlation features of varying strengths. Convergence monitoring results include information such as the convergence parameter τ, oscillation characteristics, and abnormal patterns across various dimensions. These surface manifestations conceal complex causal driving mechanisms, requiring a systematic approach to identify. Strong correlation factors represent factors that have a direct and significant impact on system performance. For example, control parameters directly determine response speed, and actuator characteristics directly influence tracking accuracy. Changes in these factors are immediately reflected in system performance. Weak correlation factors, while having minimal impact when acting alone, can have significant effects under specific conditions or when accumulated over long periods of time. For example, ambient temperature slowly affects device characteristics, and measurement noise gradually accumulates, causing control deviation. Causal strength is determined using the path coefficient β = |ΔY / ΔX| × P(X→Y), where ΔY / ΔX represents the sensitivity of Y to changes in X, and P(X→Y) represents the conditional probability that X causes Y. A β > 0.7 indicates a strong correlation, while a β < β ≤ 0.7 indicates a weak correlation. The core value of hierarchical identification lies in its ability to distinguish between primary and secondary factors, allowing the control system to prioritize critical factors while not neglecting minor but persistent influences. For example, in robotic control, excessive control gain leading to speed oscillation is identified as a strong correlation (β=0.85), requiring immediate adjustment; while steady-state offset caused by ambient temperature changes is identified as a weak correlation (β=0.42), which can be resolved through slow compensation. Through a scientific hierarchical identification method, strong and weak correlation factors are accurately obtained.
[0069] Based on strong and weak correlation factors, timing characteristics are analyzed to refine the time scale of causal transmission. Even for strongly correlated factors, some effects are almost instantaneous, such as the transition from control command to motor response; others require time to manifest, such as the transition from load change to steady-state error. This difference in time scales dictates different control strategies: fast causal relationships require real-time response and rapid compensation, while slow causal relationships can be addressed through prediction and gradual adjustment. Time delay analysis is determined using the cross-correlation function R_xy(τ)=E[(X_t-μ_x)(Y_{t+τ}-μ_y)] / (σ_xσ_y), where τ is the time delay, μ_x and μ_y are the means of X and Y, respectively, and σ_x and σ_y are the standard deviations. The τ value corresponding to the function peak is the causal delay time. Fast response is defined as τ < 0.5s, and slow response is defined as τ ≥ 0.5s. The threshold is selected based on the sampling period and dynamic response requirements of the control system. Time series analysis of strong correlation factors helps identify transient issues requiring high-frequency controllers and gradual changes that can be addressed through parameter adaptation. Time series analysis of weak correlation factors is equally important. Fast weak correlations, such as high-frequency noise, require filtering, while slow weak correlations, such as performance degradation, require trend prediction. Through time series analysis, four types of spatiotemporal correlation factors were successfully identified: fast strong correlation, slow strong correlation, fast weak correlation, and slow weak correlation.
[0070] A multi-layered feedback network with different layers and speeds is constructed based on four types of spatiotemporal correlation factors, forming a clearly defined control architecture. The core control layer, composed of fast and strong correlation factors, undertakes the system's basic control functions and must provide the fastest response and highest reliability. Any delays or failures will directly impact system performance. The steady-state regulation layer addresses slow and strong correlations. While responding more slowly, their impact is far-reaching. This layer ensures long-term system stability and avoids performance drift through continuous parameter optimization and compensation. The disturbance compensation layer specifically addresses disturbances caused by fast and weak correlations. While these disturbances may have minimal impact individually, if left unchecked, they can propagate and amplify throughout the system, ultimately impacting control quality. The trend prediction layer focuses on slow and weak correlations, predicting system evolution trends through long-term data analysis, providing a basis for preventive maintenance and performance optimization. The inter-layer connection weight is defined as w_ij = β_ij × exp(-d_ij / d_0), where β_ij is the causal strength from node i to j, d_ij is the inter-layer distance (the difference in the number of layers), and d_0 is a normalization constant (typically 1). This exponential decay reflects the rapid attenuation of cross-layer influences. Each layer has an independent update frequency: 1kHz for the core layer, 10Hz for the steady-state layer, 100Hz for the perturbation layer, and 1Hz for the prediction layer. This multi-rate design significantly improves computational efficiency. Through layered mapping, a multi-level feedback network was successfully constructed.
[0071] A prioritized path search based on a multi-layered feedback network extracts a complete causal chain. Path search is not a simple graph traversal, but rather a comprehensive optimization problem that considers multiple factors, including causal strength, time delay, and path length. Prioritization ensures that the dominant paths with the greatest system impact are identified first. These paths typically pass through the core control layer and exhibit strong causal relationships and fast transmission. The path importance score is calculated using the formula S = G × exp(-T / T_0), where G = Πβ_i is the product of all causal strengths along the path (the total path gain), T = Στ_i is the sum of all delays along the path, and T_0 is a time normalization constant (typically the system time constant). This score comprehensively considers both causal strength and time impact. After identifying the dominant path, the algorithm searches for secondary but still important auxiliary paths. These paths may pass through other layers and, while less influential, are essential for a complete understanding of the system. Preserving timing information is crucial, with the delay of each node recorded. This allows the final causal chain to describe not only "what influences what" but also "how quickly the impact occurs." The search results are integrated to remove redundant paths, merge similar branches, and retain complete chains with unique contributions. For example, a typical dominant chain: control gain K_p → (τ = 0.1s, β = 0.9) → velocity response → (τ = 0.3s, β = 0.8) → position tracking → (τ = 0.5s, β = 0.7) → convergence performance, with a total delay of 0.9s and a path gain of 0.504. Through a systematic prioritized search, a complete causal feedback chain with spatiotemporal characteristics is generated.
[0072] Predictive correction of subsequent manipulation actions is carried out using causal feedback chains to achieve feedforward compensation and performance optimization. The causal feedback chain C = {c_1 → c_2 →... → c_n} reveals the transfer mechanism from control input to performance output and can be used for prediction and improvement. The prediction model is constructed based on the chain: y(t + h) = Σα_i × u(t - d_i), where h is the prediction horizon, d_i is the causal delay, and α_i is the transfer coefficient. The correction strategy design adopts the idea of model predictive control and makes early adjustments when detecting adverse trends. The feedforward compensation applies a reverse correction u_ff = -β × e_predicted d time in advance according to the delay characteristics of the causal chain. Parameter optimization blocks adverse causal transmission by adjusting the control parameters at the source of the chain. Online learning continuously updates the causal model parameters to adapt to system changes. The correction amount calculation considers the comprehensive influence of multiple causal chains: Δu = Σw_i × Δu_i, where the weights reflect the importance of the chains. The stability constraint ensures that the correction does not introduce new problems, |Δu| < u_max × 0.2. For example, in the task of assisting the elderly in transferring between a bed and a wheelchair, causal chain analysis reveals the transmission path of "hip joint driver response delay (0.2s) → increased body tilt angular velocity (0.5s) → center of gravity offset exceeding the limit (1.0s) → possible risk of falling". Predictive corrections are immediately taken: increasing the knee joint support torque 0.2s in advance to compensate for the hip joint delay, and at the same time reducing the overall movement speed by 20% to prevent rapid center of gravity offset. After the correction, the body tilt angle of the elderly remains within the safe range of ±5°, and the smooth transfer from the wheelchair to the bed is successfully completed. Through predictive analysis and active correction based on causal feedback chains, potential problems can be identified in advance and targeted measures can be taken, achieving predictive correction of subsequent manipulation actions and completing high-quality remote control of humanoid robots.
[0073] In order to execute the remote control method of the humanoid robot based on wireless communication corresponding to the above method embodiment to achieve the corresponding functions and technical effects. Refer to Figure 2 , Figure 2 FIG. shows a structural block diagram of a remote control device 200 for a humanoid robot based on wireless communication provided by an embodiment of the present application. For the sake of convenience of description, only the parts related to this embodiment are shown. The remote control device 200 for a humanoid robot based on wireless communication provided by an embodiment of the present application includes: An instruction parsing module 201, configured to receive a remote control instruction from a manipulator, perform semantic parsing on the remote control instruction to generate manipulation intention data, establish a two-way expectation calibration channel based on the manipulation intention data, and generate an expectation consistency parameter and a deviation tolerance threshold based on the two-way expectation calibration channel; A reverse inference module 202, configured to obtain an expected execution result based on the expectation consistency parameter, reverse-infer an optimal manipulation sequence according to the expected execution result, and perform negative space analysis on the optimal manipulation sequence to identify a set of prohibited execution actions; A paradox verification module 203 is configured to construct a paradox verification scenario by combining the optimal control sequence and the prohibited action set, and to generate control boundary parameters by performing a conflict test on the paradox verification scenario; a safety control module 204, configured to establish an interlock verification condition based on the manipulation boundary parameter and the deviation tolerance threshold, and generate a safety unlocking sequence based on the interlock verification condition; An execution monitoring module 205 is configured to monitor the optimal control sequence in parallel based on the security unlocking sequence to obtain execution verification data, determine a target execution branch based on the execution verification data, and generate a unified instruction stream based on the target execution branch; The communication execution module 206 is used to transmit the unified instruction stream to the robot side via wireless communication, perform instruction verification on the robot side based on the safety unlock sequence to generate a verification pass signal, and drive the robot to execute and obtain actual execution status data based on the verification pass signal; The feedback optimization module 207 is used to perform convergence analysis based on the actual execution state data and the deviation tolerance threshold to generate a convergence monitoring result, obtain a causal feedback chain based on the convergence monitoring result, and use the causal feedback chain to predictively correct subsequent control behaviors to complete remote control.
[0074] The aforementioned wireless communication-based humanoid robot remote control device 200 can implement the wireless communication-based humanoid robot remote control method of the aforementioned method embodiment. The optional options in the aforementioned method embodiment also apply to this embodiment and will not be described in detail here. The remaining contents of the present application embodiment can be referred to the contents of the aforementioned method embodiment and will not be further described in this embodiment.
[0075] like Figure 3 As shown, the third embodiment of the present invention further provides a computer device, including a memory 301, a processor 302, and a computer program stored in the memory 301 and executable on the processor 302, characterized in that when the processor 302 executes the program, the steps of the method for remote control of a humanoid robot based on wireless communication described in the first embodiment of the present invention are implemented.
[0076] The purpose of the above embodiments is to exemplify and deduce the technical solution of the present invention, and to fully describe the technical solution, purpose and effect of the present invention. Its purpose is to enable the public to have a more thorough and comprehensive understanding of the disclosed content of the present invention, and it does not limit the scope of protection of the present invention.
[0077] The above embodiments are not exhaustive and may include many other embodiments not listed above. Any replacements and improvements made without violating the concept of the present invention are within the scope of protection of the present invention.
Claims
1. A method for remotely controlling a humanoid robot based on wireless communication, characterized in that: include: receiving a remote control command from an operator, performing semantic analysis on the remote control command to generate control intention data, establishing a two-way expectation calibration channel based on the control intention data, and generating an expected consistency parameter and a deviation tolerance threshold based on the two-way expectation calibration channel; Obtaining an expected execution result based on the expected consistency parameter, inferring an optimal control sequence based on the expected execution result, and performing negative space analysis on the optimal control sequence to identify a prohibited action set; Combining the optimal control sequence and the prohibited action set to construct a paradox verification scenario, and performing a conflict test on the paradox verification scenario to generate control boundary parameters; establishing an interlock verification condition based on the manipulation boundary parameter and the deviation tolerance threshold, and generating a safe unlocking sequence based on the interlock verification condition; Parallel monitoring of the optimal control sequence based on the security unlocking sequence to obtain execution verification data, determining a target execution branch according to the execution verification data, and generating a unified instruction stream based on the target execution branch; The unified instruction stream is transmitted to the robot side via wireless communication, the robot side performs instruction verification based on the safety unlocking sequence to generate a verification pass signal, and the robot is driven to execute and obtain actual execution status data based on the verification pass signal; A convergence analysis is performed based on the actual execution state data and the deviation tolerance threshold to generate a convergence monitoring result, a causal feedback chain is obtained according to the convergence monitoring result, and the causal feedback chain is used to predictively correct subsequent control behaviors to complete remote control.
2. The method according to claim 1, characterized in that The reverse deducing of the optimal control sequence according to the expected execution result includes: Constructing a result status mapping table based on the expected execution result; Reverse-analyze the result state mapping table to obtain a causal association path; The optimal control sequence is generated by reverse deduction based on the causal relationship path.
3. The method according to claim 1, characterized in that The performing negative space analysis on the optimal control sequence to identify a prohibited action set includes: Constructing a full action space matrix corresponding to the optimal control sequence; Excluding the optimal manipulation sequence from the full action space matrix to obtain a negative space region; Analyzing the negative space area to generate a dangerous action mark; A prohibited action set is acquired based on the dangerous action mark.
4. The method according to claim 1, wherein The step of constructing a paradox verification scenario by combining the optimal manipulation sequence and the prohibited action set includes: Performing intersection analysis on the optimal control sequence and the prohibited action set to obtain an intersection analysis result; Identifying potential conflicting nodes based on the intersection analysis results; Performing conflict reinforcement processing on the potential conflict node to generate a conflict amplification factor; The conflict amplification factor is used to construct a paradox verification scenario.
5. The method according to claim 1, wherein The generating of the control boundary parameters by performing the conflict test through the paradox verification scenario includes: simulating the execution of contradictory instructions in the paradox verification scenario; Analyze the execution process of conflicting instructions and obtain conflict response data; Performing extreme value analysis based on the conflict response data to determine a tolerance threshold; A control boundary parameter is obtained according to the tolerance threshold.
6. The method according to claim 1, characterized in that The establishing of the interlock verification condition based on the manipulation boundary parameter and the deviation tolerance threshold comprises: Performing a correlation analysis on the control boundary parameter and the deviation tolerance threshold to obtain a correlation analysis result; Constructing a multidimensional verification matrix based on the association analysis results; Performing constraint screening on the multidimensional verification matrix to obtain key constraint factors; An interlock verification condition is established according to the key constraint factors.
7. The method according to claim 1, characterized in that The step of performing parallel monitoring on the optimal control sequence based on the safety unlocking sequence to obtain execution verification data includes: Decomposing the security unlocking sequence into a plurality of monitoring checkpoints; Performing real-time state sampling on the optimal control sequence at the monitoring checkpoint to obtain sampling state data; Analyzing the sampled state data to generate a state change trend; Execution verification data is generated based on the state change trend.
8. The method according to claim 1, characterized in that The obtaining of a causal feedback chain according to the convergence monitoring result includes: Performing hierarchical causal relationship identification on the convergence monitoring results to obtain strong correlation factors and weak correlation factors; Based on the strong correlation factor and the weak correlation factor, a time series characteristic analysis is performed to subdivide the strong correlation factor into fast strong correlation and slow strong correlation, and the weak correlation factor into fast weak correlation and slow weak correlation, thereby obtaining four types of spatiotemporal correlation factors; Based on the four types of spatiotemporal correlation factors, a multi-level feedback network with different layers and speeds is constructed, in which the fast strong correlation constitutes the core control layer, the slow strong correlation constitutes the steady-state regulation layer, the fast weak correlation constitutes the disturbance compensation layer, and the slow weak correlation constitutes the trend prediction layer; A prioritized path search is performed based on the multi-level feedback network to generate a causal feedback chain with hierarchical spatiotemporal characteristics.
9. A remote control device for a humanoid robot based on wireless communication, characterized in that: include: a command parsing module, configured to receive a remote control command from an operator, perform semantic parsing on the remote control command to generate control intention data, establish a two-way desired calibration channel based on the control intention data, and generate desired consistency parameters and a deviation tolerance threshold based on the two-way desired calibration channel; a reverse reasoning module, configured to obtain an expected execution result based on the expected consistency parameter, reversely infer an optimal control sequence based on the expected execution result, and perform negative space analysis on the optimal control sequence to identify a set of prohibited actions; A paradox verification module, configured to construct a paradox verification scenario by combining the optimal control sequence and the prohibited action set, and to generate control boundary parameters by performing a conflict test on the paradox verification scenario; a safety control module, configured to establish an interlock verification condition based on the manipulation boundary parameter and the deviation tolerance threshold, and generate a safety unlocking sequence based on the interlock verification condition; an execution monitoring module, configured to monitor the optimal control sequence in parallel based on the security unlocking sequence to obtain execution verification data, determine a target execution branch based on the execution verification data, and generate a unified instruction stream based on the target execution branch; a communication execution module, configured to transmit the unified instruction stream to the robot side via wireless communication, perform instruction verification on the robot side based on the safety unlock sequence to generate a verification pass signal, and drive the robot to execute and obtain actual execution status data based on the verification pass signal; A feedback optimization module is used to perform convergence analysis based on the actual execution state data and the deviation tolerance threshold to generate a convergence monitoring result, obtain a causal feedback chain based on the convergence monitoring result, and use the causal feedback chain to predictively correct subsequent control behaviors to complete remote control.
10. A computer device, characterized in that: The method comprises a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and implement the method according to any one of claims 1 to 8 when executing the computer program.
Citation Information
Patent Citations
Method for controlling and operating a production cell, and control device
CN101128306A
Master-slave surgical robot control system and method capable of suppressing tremor
CN115607297A
Intelligent remote controller control method, device and equipment based on gesture recognition
CN120386456A
Methods and systems for distributing remote assistance to facilitate robotic object manipulation
US9486921B1
Method and system for remotely monitoring and forecasting the state of technical equipment
WO2022010377A1
Cited By
Imitation learning method, device and equipment for intelligent agent with body
CN121189522A