Interaction control method and device based on multi-source perception, equipment and medium

By preprocessing and feature association of multi-source sensing data, and combining it with task objectives to generate action strategies, and adjusting control output in real time, the problem of disconnect between robot action execution and interaction state in complex environments is solved, thereby improving safety and adaptability.

CN121893232APending Publication Date: 2026-04-21PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2026-03-03
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In existing technologies, the motion planning and control of robots in complex interactive scenarios suffer from several problems: multi-source perception data is difficult to represent in a unified manner; motion strategies and execution feedback lack consistent alignment; closed-loop adjustment of execution deviations is insufficient; and safety monitoring and response rely heavily on single-point triggers, resulting in a disconnect between strategy output and actual execution, and insufficient safety and adaptability.

Method used

Collect multi-source heterogeneous sensing data from the working environment, preprocess it to generate a standardized sensing dataset, analyze the state of interactive objects and environmental constraints through feature extraction and semantic association, generate action strategies in combination with task objectives, collect execution feedback data in real time, adjust control output, perform safety monitoring and model updates, and build a closed-loop control mechanism of perception, decision-making, execution, feedback and model updates.

Benefits of technology

It enables robots to adaptively adjust their action strategies in complex environments, improving interaction safety and environmental adaptability, and ensuring consistency and stability in the execution process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121893232A_ABST
    Figure CN121893232A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent decision making, can be applied to business scenes such as financial science and technology and medical health, and discloses an interaction control method, device and equipment based on multi-source perception, and a medium, and the method comprises the steps: collecting and preprocessing multi-source heterogeneous perception data in an operation environment, and generating a standardized perception data set; analyzing the current state information of the interaction object and the constraint condition of the working environment based on the standardized sensing data; generating an action strategy by using the optimization model in combination with the task target and controlling the robot to execute; and collecting execution feedback data in the execution process, adjusting and controlling output according to the execution deviation, and performing safety monitoring and optimization model updating based on the execution feedback data. According to the method, the execution feedback data is introduced into the safety monitoring and optimization model updating process, closed-loop self-adaptive control is formed, the robot can dynamically adjust the action strategy according to the state change, and therefore the execution safety and adaptability in the complex environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent decision-making technology, and in particular to an interactive control method, device, equipment and medium based on multi-source perception. Background Technology

[0002] In existing technologies, robot motion planning and control for complex interactive scenarios often presents a fragmented approach. Multi-source perception data is difficult to form a unified representation and stably support state analysis and constraint recognition. There is a lack of consistent alignment and comparison criteria between the generated motion strategy and the execution feedback. The closed-loop adjustment capability for execution deviations during execution is insufficient. Safety monitoring and response mostly rely on single-point triggering and lack continuous coupled monitoring based on constraints and execution feedback data. At the same time, execution feedback data is difficult to be precipitated into effect analysis information that can be used for model updates. Therefore, under dynamic environments and conditions with significant individual differences, problems such as disconnect between strategy output and actual execution, delayed risk response, and difficulty in continuous adaptive optimization can easily occur.

[0003] In the healthcare sector, embodied nursing robots need to perform delicate movements in scenarios characterized by sensitive contact pressure, frequent posture changes, and complex environmental constraints. Existing technologies often acquire geometric information through single sensing sources or shallow fusion, making it difficult to simultaneously characterize contact and force states. This results in insufficient analysis of the current state of the interacting object and the constraints of the working environment. Simultaneously, motion strategies are often generated based on preset templates or static parameters. While execution feedback data is collected during the execution phase, there is a lack of stable mechanisms for judging and continuously adjusting execution deviations between the feedback data and control parameters and motion paths, easily leading to problems such as trajectory deviations and overlapping contact pressure fluctuations. Regarding safety monitoring and response, common practices focus on triggering actions by exceeding the limits of a single physical quantity, lacking continuous monitoring and tiered response logic based on constraints and execution feedback data. This makes it difficult to maintain stable safety boundaries for nursing actions under dynamic interaction conditions, exacerbating the problems of process interruptions and uneven risk management.

[0004] In the fintech business, automated execution systems, when faced with dynamic business constraints and multi-source data inputs, also suffer from weak connections between data fusion, state parsing, constraint identification, strategy generation, and execution monitoring. Existing technologies often separate strategy output from execution feedback, lacking a mechanism to consistently compare execution feedback data with strategy parameters and establish quantifiable deviation metrics. This makes it difficult to identify and correct execution deviations in a timely manner. Security monitoring and response typically rely on static rules or single-point threshold triggers, making it difficult to maintain continuous monitoring and tiered response stability when constraints change. Furthermore, execution feedback data is mostly used for audit records or post-event analysis, lacking a structured path for generating effect analysis information. This makes it difficult for models or strategies to continuously update, leading to insufficient adaptability and decreased stability under business fluctuations and environmental changes. Summary of the Invention

[0005] The main objective of this invention is to provide an interactive control method, device, equipment, and storage medium based on multi-source perception, aiming to solve the technical problem that the lack of a unified adaptive control closed loop based on multi-source heterogeneous perception in the prior art leads to the disconnect between the robot's action execution and interactive state in dynamic environments, resulting in insufficient safety and adaptability.

[0006] To achieve the above objectives, the present invention provides an interactive control method based on multi-source sensing, comprising: Collect multi-source heterogeneous sensing data from the working environment, and preprocess the multi-source heterogeneous sensing data to generate a standardized sensing dataset; Feature extraction and semantic association are performed on the standardized perception dataset to obtain feature association information, and the current state information of the interactive object and the constraints of the working environment are parsed based on the feature association information. Based on the current state information and the constraints, and combined with the task objective, an optimization model is used to solve the problem and generate an action strategy that includes control parameters and motion path. The robot's execution device is controlled to operate according to the aforementioned action strategy, and execution feedback data is collected in real time during the operation. Determine the execution deviation between the execution feedback data and the control parameters and motion path in the action strategy, and adjust the control output of the robot execution device according to the execution deviation; Based on the constraints and the execution feedback data, security monitoring and response are performed, and effect analysis information is generated based on the execution feedback data to update the optimization model.

[0007] Furthermore, to achieve the above objectives, the present invention provides an interactive control device based on multi-source sensing, comprising: The multi-source sensing data fusion module is used to collect multi-source heterogeneous sensing data in the working environment and preprocess the multi-source heterogeneous sensing data to generate a standardized sensing dataset. The state and constraint parsing module is used to extract features and semantically associate them with the standardized perception dataset to obtain feature association information, and to parse the current state information of the interactive object and the constraints of the working environment based on the feature association information. The action strategy generation module is used to generate an action strategy containing control parameters and motion path by solving the current state information and the constraints, combined with the task objective, using an optimization model. The execution and feedback acquisition module is used to control the robot execution device to operate according to the action strategy and to collect execution feedback data in real time during the operation. An execution deviation adjustment module is used to determine the execution deviation between the execution feedback data and the control parameters and motion path in the action strategy, and adjust the control output of the robot execution device according to the execution deviation; The safety monitoring and model update module is used to perform safety monitoring and response based on the constraints and the execution feedback data, and to generate effect analysis information based on the execution feedback data to update the optimization model.

[0008] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a multi-source perception-based interactive control program stored in the memory and executable on the processor, wherein the multi-source perception-based interactive control program, when executed by the processor, implements the steps of the multi-source perception-based interactive control method as described above.

[0009] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a multi-source perception-based interactive control program, wherein the multi-source perception-based interactive control program, when executed by a processor, implements the steps of the multi-source perception-based interactive control method as described above.

[0010] Beneficial Effects: This invention relates to the field of intelligent decision-making technology and can be applied to business scenarios such as fintech and healthcare. It discloses an interactive control method, device, equipment, and medium based on multi-source perception, comprising: collecting multi-source heterogeneous perception data from the working environment and preprocessing it to generate a standardized perception dataset; performing feature extraction and semantic association based on the standardized perception dataset to parse the current state information of the interactive object and the constraints of the working environment; generating an action strategy including control parameters and motion paths using an optimization model in conjunction with the task objective; controlling the robot's execution device to operate according to the action strategy and collecting execution feedback data; adjusting the control output based on the execution deviation between the execution feedback data and the action strategy; performing safety monitoring and response based on constraints and execution feedback data, and generating effect analysis information for optimizing model updates. This invention constructs a closed-loop control mechanism of perception, decision-making, execution, feedback, and model updates, enabling execution feedback data to continuously influence safety monitoring and optimize model updates, thereby achieving adaptive adjustment of the action strategy and improving the robot's interactive safety and environmental adaptability in complex environments. Attached Figure Description

[0011] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a schematic diagram of an application environment for an interactive control method based on multi-source sensing in one embodiment of the present invention; Figure 2 This is a flowchart illustrating an embodiment of the interactive control method based on multi-source sensing of the present invention. Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the interactive control device based on multi-source sensing of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0012] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0013] The interactive control method based on multi-source sensing provided in this invention can be applied to applications such as... Figure 1In this application environment, the client communicates with the server via a network. The server can collect multi-source heterogeneous sensing data from the working environment through the client and preprocess it to generate a standardized sensing dataset. Based on the standardized sensing dataset, feature extraction and semantic association are performed to parse the current state information of the interactive object and the constraints of the working environment. Combined with the task objective, an optimization model is used to generate an action strategy containing control parameters and motion paths. The robot's execution device is controlled to run according to the action strategy and execution feedback data is collected. The control output is adjusted according to the execution deviation between the execution feedback data and the action strategy. Based on the constraints and execution feedback data, safety monitoring and response are performed, and effect analysis information is generated for optimizing model updates. This invention constructs a closed-loop control mechanism of perception, decision-making, execution, feedback, and model updates, so that execution feedback data continuously affects safety monitoring and optimizing model updates, thereby achieving adaptive adjustment of the action strategy and improving the robot's interaction safety and environmental adaptability in complex environments. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster composed of multiple servers. The invention will be described in detail below through specific embodiments.

[0014] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the interactive control method based on multi-source perception provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0015] like Figure 2 As shown, the interactive control method based on multi-source perception proposed in this invention includes the following steps: S10: Collect multi-source heterogeneous sensing data in the working environment, and preprocess the multi-source heterogeneous sensing data to generate a standardized sensing dataset. In this embodiment, the problem to be solved is that the multi-source heterogeneous sensing data in the working environment has inconsistent sampling rhythms, data formats, and units and coordinate references. Directly entering it into subsequent processing will introduce alignment errors and semantic ambiguities. Therefore, the original input is converted into a unified and computable data set. Acquisition corresponds to obtaining multi-source heterogeneous sensing data from the working environment. Specifically, the controller initiates synchronous or near-synchronous sampling to multiple sensing channels. The sampling record carries at least the channel identifier and sampling time mark to ensure that different sources can be distinguished and traced. Multi-source corresponds to more than one sensing channel, which can come from tactile, visual, force, distance, inertial, or state quantity channels, etc. Heterogeneous corresponds to different channels that differ in dimension, sampling rate, and data type. The organization of sensing data requires establishing field definitions for each channel, specifying the unit, value range, and missing value marker. Preprocessing transforms the raw records into an aligned and comparable intermediate form. This includes timestamp alignment, noise suppression, missing data handling, unit conversion, and coordinate unification. Timestamp alignment resamples each channel to a common time grid; noise suppression addresses spikes and high-frequency jitter; missing data handling uses masking or interpolation to fill in missing segments; unit conversion maps sensor outputs to a unified unit; and coordinate unification maps spatially relevant quantities to a unified reference frame. Generation involves summarizing the preprocessed results into a standardized sensing dataset. Standardization emphasizes unified field naming, units, time references, and coordinate references. The dataset includes both the data ontology and necessary explanatory information, enabling subsequent processing to directly read and correctly interpret the meaning of each field.

[0016] This embodiment performs timestamp alignment, noise suppression, missing data processing, unit conversion, and coordinate unification on multi-source heterogeneous sensing data to generate a standardized sensing dataset. Data from different channels form a consistent expression in terms of time, unit, and spatial reference, reducing fluctuations caused by incomparability and interpretation ambiguity of cross-channel data, and enabling the input data to have the attributes of being alignable, comparable, and reusable.

[0017] S20, perform feature extraction and semantic association on the standardized perception dataset to obtain feature association information, and parse the current state information of the interactive object and the constraints of the working environment based on the feature association information. In this embodiment, the problem to be solved is that although the standardized sensing dataset has achieved unification of time, units, and coordinate benchmarks, different channels still exist in the form of discrete fields, lacking cross-channel correspondences, making it difficult to directly use for judging the state of interactive objects and constraints of the working environment. Therefore, the standardized sensing dataset is converted into feature association information that can express a relational structure, and on this basis, the current state information of the interactive object and the constraints of the working environment are output. Feature extraction of the standardized sensing dataset corresponds to extracting feature quantities that can represent morphology, contact, change trends, and abnormal fluctuations from the dataset. In specific implementation, continuous quantities, sequential quantities, and spatial quantities are first distinguished based on field type. Statistical quantities and change quantities are calculated for continuous quantities, segment representations are calculated for sequential quantities, and geometric representations and relative relationship representations are calculated for spatial quantities. The feature quantities and the original fields maintain a traceable mapping to avoid the inability to locate the source in subsequent parsing. Semantic association establishes relationships between features from different sources under the same semantic object. The semantic object can correspond to parts of the interactive object, interactive areas, passable areas, obstacle entities, or contact events. In practice, features within a time window are aligned, and association edges are established based on spatial proximity, temporal co-occurrence, and consistent changes. These edges are then assigned association strength or confidence, forming a structured representation usable for inference. The output of the feature association information contains structured data including nodes, edges, and attributes. Nodes are features or semantic objects, edges are association relationships, and attributes include association strength, time window identifiers, and spatial reference identifiers, enabling subsequent analysis to structurally locate the evidence chain. Based on the current state information of the interactive object obtained from the feature association information, aggregated inferences are performed on related semantic objects. The current state information at least covers dimensions such as tolerability, stability, and risk tendency. In practice, a set of nodes related to the interactive object is selected from the feature association information, and aggregated quantities such as contact intensity distribution, posture change amplitude, and state fluctuation amplitude are calculated. Discrete levels or continuous scores are output according to preset mapping rules or learnable mappings, while retaining the triggering criteria to support consistency checks. Based on the feature association information parsing, the constraints of the working environment are obtained and the semantic objects related to the working environment are restricted and extracted. The constraints should at least cover categories such as the boundary of the passable area, the safe distance of obstacles, and the accessibility area restriction. In specific implementation, environmental entities and spatial relationship nodes are screened from the feature association information, the distribution of obstacles and the range of free space are derived, geometric boundary expressions and distance threshold expressions are generated, and the constraints are written into structural fields that can be directly used for subsequent solutions.

[0018] This embodiment extracts traceable features from a standardized perception dataset and establishes cross-channel semantic associations to obtain feature association information. Then, it outputs the current state information of the interactive object and the constraints of the working environment from the feature association information. The input data is transformed from a set of fields into a structured association expression, reducing misjudgments caused by the fragmentation of cross-channel information, and making the state output and the environmental constraint output have an interpretable source link and a consistent expression.

[0019] S30, Based on the current state information and the constraints, and in conjunction with the task objective, an optimization model is used to solve the problem and generate an action strategy that includes control parameters and motion path. In this embodiment, the current state information and constraints are used to define the tolerable range of the interactive object and the traversable range of the working environment. However, there is a lack of output that can directly drive the robot's execution device. Therefore, the current state information, constraints, and task objective are input into the optimization model together to complete the solution and generate the action strategy. The current state information comes from the analysis results of the previous processing stage and is used to form the state input quantity and constrain the control strength and attitude change amplitude. The constraints come from the analysis results of the previous processing stage and are used to form the feasible domain expression and limit the spatial range of the motion path and obstacle avoidance requirements. The task objective is used to define the desired final state and process preference, which can be transformed into the terminal condition and process cost of the objective function. The optimization model is used to uniformly organize the objective function, constraints, and decision variables. The decision variables are used to characterize the action output to be solved and can be the parameterized representation of the motion path and the set of values ​​of the control parameters. The solution process involves searching for solutions that satisfy the constraints and meet the task objectives within the feasible region. After the solution is obtained, it is organized into an action strategy. The motion path describes the spatial evolution of the robot's actuator, and the control parameters describe the control requirements when executing along the motion path. The action strategy aligns the motion path and control parameters with the same time reference so that they can be directly issued for execution.

[0020] The decision variables for the optimization model can be either joint space parameters or end-effector space parameters; constraints can be expressed using inequality constraints or distance constraints; the task objective can be expressed in a single-objective form or a weighted multi-objective form; the solution process can employ iterative search or sampling-based search; control parameters can be output as a combination of contact force-related parameters, velocity-related parameters, or acceleration-related parameters, and aligned with the parameterized curve of the motion path on the time axis to form an action strategy.

[0021] This embodiment incorporates the current state information, constraints, and task objectives into the optimization model for solution and outputs an action strategy that includes both control parameters and motion paths. This allows the action output to be directly executed within the feasible spatial domain and under state constraints, while maintaining consistency with the task objective.

[0022] S40, control the robot execution device to operate according to the action strategy, and collect execution feedback data in real time during the operation; In this embodiment, the motion strategy, as the output of the previous processing stage, includes two categories: motion path and control parameters. It is used to directly drive the robot's actuator to produce physical actions. Controlling the robot actuator to operate according to the motion strategy means mapping the motion path in the motion strategy to continuous pose change commands, and mapping the control parameters to control quantities such as force, velocity, or acceleration, which are then sent to the actuator's drive unit through a control interface. The running process corresponds to the actuator generating continuous actions in physical space. This process is bound to the time axis to ensure that path advancement and control quantity changes are synchronized. During operation, execution feedback data is collected in real time to reflect the actual interaction state between the actuator and the interactive object. The execution feedback data originates from the actuator's internal state perception and external interaction perception, including the actuator's current output state, the actual force state at the end effector, and the spatial pose change state. Real-time acquisition means continuously updating the feedback data while the action is running, ensuring that the feedback data corresponds one-to-one with the action advancement process and maintains temporal consistency.

[0023] The motion path in the motion strategy can be converted into a joint angle sequence and sent to the underlying drive controller, or it can be converted into an end-effector pose trajectory and further processed by the inverse kinematics module. Control parameters can be sent synchronously with the trajectory as independent control variables, or they can be embedded in the trajectory point attributes and sent uniformly. Execution feedback data can be collected through the force sensing unit to acquire the contact state, through the position encoding unit to acquire the pose state, or through the controller's internal state interface to acquire the drive output state, and then aligned and cached on the time axis.

[0024] This embodiment continuously acquires execution feedback data during the operation, enabling the actual operating state of the execution device to be fully recorded and kept synchronized with the action strategy, providing a stable data foundation for subsequent analysis and adjustment based on the operating state.

[0025] S50, determine the execution deviation between the execution feedback data and the control parameters and motion path in the action strategy, and adjust the control output of the robot execution device according to the execution deviation; In this embodiment, the execution feedback data reflects the actual behavior of the actuator in the real operating environment, while the control parameters and motion path in the motion strategy reflect the expected behavior. Determining the execution deviation between the execution feedback data and the control parameters and motion path represents aligning and comparing the actual behavior with the expected behavior under the same time reference. This comparison process operates simultaneously on both the control parameter dimension and the spatial path dimension. In the control parameter dimension, the parameter deviation is formed by comparing the actual output quantity with the target control quantity; in the motion path dimension, the spatial offset is formed by comparing the actual pose trajectory with the target motion path. The execution deviation is used to characterize the degree of deviation of the actuator's current behavior from the motion strategy and serves as the basis for subsequent control output adjustments. Adjusting the control output of the robot actuator based on the execution deviation means mapping the execution deviation to a correction amount for the control output and updating the control commands within the current control cycle, causing the actuator's operating state to converge towards the target state constrained by the motion strategy, thereby forming a continuous closed-loop adjustment relationship.

[0026] Execution deviation can be calculated by subtracting the target parameter from the actual feedback data, or it can be comprehensively expressed by an error function to represent the multidimensional deviation. Control output adjustment can be applied directly to the original control command by superimposing correction values, or it can be applied indirectly to the actuator by adjusting the control gain or limiting the output variation range. In different operating scenarios, emphasis can be placed on path deviation correction or control parameter deviation correction to adapt to different requirements for interaction accuracy and stability.

[0027] This embodiment continuously identifies execution deviations and synchronously adjusts the control output, enabling the execution device to dynamically correct deviations during operation, thereby improving the consistency and stability between action execution and action strategy.

[0028] S60, perform safety monitoring and response based on the constraints and the execution feedback data, and generate effect analysis information based on the execution feedback data to update the optimization model.

[0029] In this embodiment, constraints are used to limit the permissible range of behavior of the execution device in the operating environment, and execution feedback data is used to characterize the actual state changes of the execution device during operation. Safety monitoring and response based on constraints and execution feedback data means that during execution, the execution feedback data is continuously mapped to the safe range defined by the constraints to determine whether the execution state is within the permissible range. When the execution feedback data touches or approaches the limit boundary corresponding to the constraint, a response behavior matching the current risk level is triggered, keeping the execution process under control within the constraints.

[0030] Based on execution feedback data, performance analysis information is generated. This involves statistically analyzing, aggregating, and summarizing the features of the feedback data generated during execution to characterize the overall performance of the action execution in terms of stability, smoothness, and state recovery. The performance analysis information reflects the comprehensive performance of the action strategy under real execution conditions and serves as input for updating the optimization model, enabling the model to incorporate behavioral characteristics and deviation trends reflected in historical execution results during subsequent solution processes.

[0031] Safety monitoring can be achieved by comparing execution feedback data with multi-dimensional constraint thresholds in real time, or by constructing a safety status determination function to continuously evaluate the execution status. Response methods can manifest as limiting, suppressing, or terminating control outputs, or by adjusting operating modes to reduce risk levels. Effectiveness analysis information can be generated by statistically analyzing the frequency of anomalies, fluctuation amplitudes, and recovery times during execution, or by constructing comprehensive scoring indicators to express execution quality and drive the direction and magnitude of parameter updates in the optimization model.

[0032] This embodiment combines security monitoring with execution effect analysis, enabling the execution process to accumulate feedback information that can be used for model updates while meeting constraints, thereby improving the optimization model's adaptability to the actual execution environment.

[0033] In one embodiment, step S10 above includes: S101, control the contact sensing unit set at the operating end of the robot execution device to collect contact sensing data, control the vision acquisition unit to collect visual sensing data of the working environment, and obtain the internal state data of the interactive object through the data interface, and construct multi-source heterogeneous sensing data based on the contact sensing data, the visual sensing data and the internal state data. S102, using the hardware trigger signal as a unified time reference, perform timestamp alignment on the contact perception data, the visual perception data and the internal state data in the multi-source heterogeneous perception data; S103, Perform filtering processing on the timestamp-aligned contact perception data, visual perception data and internal state data respectively to obtain filtered contact perception data, filtered visual perception data and filtered internal state data. S104, the filtered contact perception data, the filtered visual perception data, and the filtered intrinsic state data are mapped to a unified spatial coordinate system for data format normalization, generating a standardized perception dataset.

[0034] In this embodiment, the contact sensing unit is used to form raw observations related to interactive contact at the end effector of the robot. Contact sensing data is used to characterize the force distribution, contact duration, or changes in contact events during the contact process. The acquisition action is implemented by configuring the sampling frequency, sampling window, and sampling trigger conditions of the contact sensing unit, ensuring that the contact sensing data corresponds to the actual interactive behavior of the end effector in both spatial location and time sequence. The visual acquisition unit is used to form visual observations of the working environment. Visual sensing data is used to characterize the appearance structure, boundary contours, relative positions, and traversable areas of the interactive object and the working environment. The acquisition action is implemented by setting exposure parameters, frame rate, resolution, and field of view, ensuring the visual sensing data has continuity and trackability in dynamic scenes. The data interface is used to acquire the intrinsic state data of the interactive object. Intrinsic state data characterizes the state changes, trends, or fluctuations of the interactive object. The acquisition action is implemented by configuring the communication protocol, data field mapping, and acquisition cycle of the data interface, ensuring that the intrinsic state data maintains consistency with the state update rhythm of the interactive object. Multi-source heterogeneous sensing data is composed of contact sensing data, visual sensing data, and intrinsic state data. This means that data from different sources, modalities, and sampling mechanisms are organized into a unified multi-source set within the same acquisition cycle, enabling subsequent processing to simultaneously utilize contact information, visual information, and intrinsic state information to complete joint analysis within the same time slice.

[0035] Hardware trigger signals serve as a unified time base to establish a global time reference for multi-source heterogeneous sensing data. Timestamp alignment operations are used to eliminate time offsets caused by differences in sampling times, transmission delays, and buffer jitter among contact sensing data, visual sensing data, and intrinsic state data. Timestamp alignment can be achieved by writing common-source timestamps to each data stream, offset compensation for arrival times, interpolation reconstruction of missing segments, or pruning and resampling of redundant segments. This ensures that contact sensing data, visual sensing data, and intrinsic state data form an aligned sample set under the same timestamp index, thereby giving contact observations, visual observations, and intrinsic state observations corresponding to the same timestamp a consistent semantic orientation.

[0036] Filtering is used to suppress noise and abnormal fluctuations in timestamp-aligned contact-sensing data, visual-sensing data, and intrinsic state data, preventing amplification by short-term spikes, random disturbances, or communication jitter during subsequent fusion. Filtering of contact-sensing data can perform smoothing, amplitude limiting, or detrending processing on high-frequency spikes, sensor drift, and instantaneous jitter in the contact pressure sequence, ensuring the contact features maintain discernible continuity under changing force. Filtering of visual-sensing data can perform denoising, stabilization, or key region enhancement processing on image noise, motion blur, or inter-frame jitter, ensuring the visual structure remains trackable in dynamic scenes. Filtering of intrinsic state data can perform smoothing, completion, or consistency verification processing on sampling jitter, missing measurements, and abnormal jumps in the state fluctuation sequence, ensuring the intrinsic state data has usability at the trend level. The results of the filtering process form filtered contact-sensing data, filtered visual-sensing data, and filtered intrinsic state data, representing the suppression of noise and abnormal components to an acceptable range for subsequent processing without altering the data source and field semantics.

[0037] Filtered contact-sensing data, filtered visual-sensing data, and filtered intrinsic state data are mapped to a unified spatial coordinate system for data format normalization. This represents a consistent processing of the spatial and data representations of multi-source data, enabling the joint use of different modal data within the same spatial reference frame. The spatial coordinate system serves as a unified reference frame, and the mapping action establishes the spatial correspondence between the various data. Contact-sensing data can map contact sampling points to the spatial coordinate system through the geometric model of the operating end or the sensor mounting pose. Visual-sensing data can map the target position in visual observation to the spatial coordinate system through camera calibration parameters and extrinsic parameter relationships. Intrinsic state data can associate state quantities with corresponding entities in the spatial coordinate system through binding relationships with interactive object identifiers or part identifiers. Data format normalization unifies the data in terms of data type, dimension, field structure, and encoding method. For example, discrete sampling sequences are converted into time slice vectors of uniform length, visual representations of different resolutions are converted into feature expressions of uniform scale, and the field set of intrinsic state data is mapped to a fixed field order and data type set, ensuring consistency in storage structure and access methods for cross-source data. The standardized sensing dataset is formed by the above mapping and normalization results, and is used to provide a data input basis that is consistent in time, space and format in subsequent processing stages.

[0038] For example, in the process of medical assistive robot care, the real-time synchronous acquisition of tactile and visual perception information is crucial to ensuring accurate motion planning and efficient care. To address the problem of single-source and inefficient fusion of perception information in existing technologies, this embodiment utilizes a multi-source sensor collaborative configuration to collect tactile, visual, and patient physiological information at the end effector of the robot and in the care environment. This ensures high-quality, accurate, and timely data, thereby providing comprehensive support for subsequent state analysis and motion planning.

[0039] In terms of tactile information acquisition, arrayed tactile sensors are integrated into the operating ends of robotic actuators (such as mechanical grippers or assisted arms) and the surfaces of care components (such as turning mats or transfer supports). These sensors have a sampling frequency of 500Hz, a measurement range of 0 to 50N, and an accuracy of ±0.01N, enabling efficient acquisition of information such as patient skin pressure distribution and limb contact stiffness. This tactile sensor can effectively measure the contact force on the patient's body surface, thereby accurately reflecting the patient's current posture, pressure status, and potential care risks.

[0040] In terms of visual information acquisition, high-definition medical cameras and depth cameras are used to collect data such as patient limb posture, facial expressions, and wound location. The high-definition medical camera has a resolution of 2560×1440 and a sampling frequency of 30Hz, enabling accurate recording of the patient's dynamic changes. The depth camera is used to acquire three-dimensional structural information of the nursing scene, with a measurement range of 0.3-3m and an accuracy of ±1mm. This information helps the robot understand the patient's precise location and the surrounding environment, providing fundamental data for motion planning and real-time feedback.

[0041] Regarding the collection of patient physiological information, a wireless vital signs monitoring module synchronously acquires patients' physiological data in real time, such as heart rate, blood pressure, and blood oxygen saturation, at a sampling frequency of 1Hz. The collection of physiological data helps assess the patient's physiological state and adaptability to nursing procedures, further ensuring the comfort and safety of nursing operations.

[0042] The data synchronization preprocessing stage employs hardware triggering to ensure time synchronization of multi-source data, with a maximum synchronization error not exceeding 1ms. Tactile signals are processed using a Butterworth low-pass filter with a cutoff frequency set to 10Hz to eliminate high-frequency noise; visual images undergo Gaussian filtering to eliminate interference from ambient lighting changes; and physiological data are processed using a Kalman filter to suppress signal drift. This preprocessing workflow ensures the stability and reliability of the acquired data.

[0043] After preprocessing, tactile, visual, and physiological data are converted into standardized datasets and provided to the robot for subsequent perception fusion, state analysis, and action decision-making. Through this multi-source information synchronous acquisition process, the robot can gain a comprehensive understanding of the patient and the care environment, thereby improving the accuracy and safety of the care process.

[0044] This embodiment constructs multi-source heterogeneous sensing data composed of contact sensing data, visual sensing data, and intrinsic state data. It uses hardware trigger signals to complete timestamp alignment, performs filtering processing on each data source, and normalizes the data format in a unified spatial coordinate system to form a standardized sensing dataset. This ensures consistency of multi-source data in the time, space, and data representation dimensions, thereby reducing the risk of mismatch when using multimodal information in combination and improving the stability and usability of subsequent processing.

[0045] In one embodiment, step S20 above includes: S201, extract skeleton posture features from the visual perception data contained in the standardized perception dataset, extract pressure distribution features from the contact perception data contained in the standardized perception dataset, and extract fluctuation features from the intrinsic state data contained in the standardized perception dataset. S202, semantically fuse the skeleton posture features, the pressure distribution features, and the wave features to generate feature association information; S203, based on the feature association information, analyze the physical vulnerability level and real-time stability of the interactive object, and generate the current state information of the interactive object; S204, based on the feature association information, identify the distribution of obstacles and passable space in the work environment, and based on the obstacle distribution, passable space and the current state information of the interactive object, generate the constraints of the work environment.

[0046] In this embodiment, the standardized perception dataset serves as a unified input carrier, containing three types of data sources: visual perception data, contact perception data, and intrinsic state data. Its organization has already achieved temporal, spatial, and format consistency. Therefore, the feature extraction stage can directly decouple and extract features based on the target semantics, aligning and combining them under the same indexing system. Visual perception data provides information on the appearance structure and spatial configuration of interactive objects. Skeleton posture features characterize the set of key points, key point connections, and the pose state of key points in the spatial coordinate system. The extraction process can be implemented based on key point detection networks, pose regression networks, or geometric constraint fitting, enabling skeleton posture features to simultaneously possess dimensions such as location, joint angles, and overall orientation, thereby supporting the characterization of body shape changes and posture shifts. Contact perception data provides information on interactive contact states and force distribution. Pressure distribution features characterize the pressure intensity, pressure gradient, contact area, and pressure center migration of the contact area between the operating end and the interactive object. The extraction process can be implemented based on sensor array reconstruction, pressure map mapping, or temporal statistics, enabling pressure distribution features to characterize different force forms such as continuous pressure, instantaneous impact, and local spikes. Intrinsic state data provides information on the state fluctuations of interactive objects. Fluctuation features characterize the amplitude, rate of change, periodicity, and abrupt changes of the intrinsic state data over time. The extraction process can be based on sliding window statistics, differential sequence analysis, or frequency domain transformation, enabling the fluctuation features to describe the degree of state stability and short-term anomalous disturbances. The three types of features are obtained from different data sources, but they maintain the same time slice and spatial semantic orientation under the common index of the standardized perceptual dataset, providing a direct correspondence for subsequent semantic-level fusion.

[0047] Semantic-level fusion is used to establish semantic relationships between skeleton posture features, pressure distribution features, and wave features, enabling multi-source features to form a reasonable relational structure rather than simply being superimposed at the splicing level. The fusion process can be organized around three stages: feature alignment, feature interaction, and association modeling. Feature alignment maps the three types of features to a unified feature space and a unified scale system, which can be achieved through linear projection, normalization, and time-slice aggregation, compressing the dimensional and distribution differences between different modalities to a comparable range. Feature interaction explicitly models intermodal dependencies, which can be achieved through attention mechanisms, gating mechanisms, or bilinear interactions, enabling skeleton posture features to establish a force correspondence between pressure distribution features and body parts, and wave features to establish a synchronous relationship between state changes and posture changes with skeleton posture features. Association modeling forms an intermediate representation that can be used for subsequent analysis. Feature association information can be a fusion vector, association matrix, graph structure data, or weighted connection set, and its content should at least include the correspondence strength, association direction, and association confidence between modalities, thus providing a unified basis for subsequent state judgment and environment analysis.

[0048] The current state information of the interactive object is obtained by parsing feature association information, and the parsing process is organized around two output dimensions. The physical vulnerability level characterizes the interactive object's tolerance to changes in force and pose. The parsing process can establish a comprehensive judgment based on the peak value and gradient of pressure distribution features, the joint constraint pattern of skeleton posture features, and the abnormal amplitude of fluctuation features, enabling the physical vulnerability level to reflect differences such as fragility, sensitivity, or tolerance. Real-time stability characterizes the controllability and fluctuation risk of the interactive object's state in the current time slice. The parsing process can establish a stability metric based on the rate of change and density of abrupt change points in fluctuation features, the micro-motion frequency of skeleton posture features, and the force drift of pressure distribution features, enabling real-time stability to reflect states such as stable, slightly fluctuating, or significantly fluctuating. The current state information is output in structured field form, which can include the physical vulnerability level, real-time stability, and confidence scores or values ​​related to both, for direct use in subsequent stages without repeated calculations.

[0049] The constraints of the working environment are obtained through semantic parsing of the environment supported by feature association information. Obstacle distribution represents the set of locations and boundaries of inaccessible, risky, and dynamically interfering areas in the working environment. The identification process can be achieved based on target detection, semantic segmentation, or depth reconstruction using visual perception data, combined with skeleton posture features to determine the space occupied and accessible areas of interactive objects, thus distinguishing between environmental obstacles and the entity boundaries of interactive objects. The traversable space represents the set of free and reachable spaces where the robot's actuators can move. The generation process can perform spatial rasterization, reachability propagation, or channel width constraint filtering based on the obstacle distribution, ensuring that the traversable space meets the actuator's geometric dimensions, kinematic constraints, and safety boundary requirements. The constraint generation action combines obstacle distribution, traversable space, and current state information to form a constraint expression oriented towards execution control. Constraints can include spatial boundaries, minimum safe distances, permissible contact areas, force restriction ranges, and posture change restriction ranges, reflecting both environmental geometric constraints and operational limitations caused by the state of interactive objects.

[0050] For example, in the healthcare field, the interaction object is the person receiving care. The standardized perception dataset includes visual perception data, contact perception data, and internal state data. Visual perception data enters the object detection network and outputs a set of key skeletal points. This set is used to calculate limb movement angles and form posture features, while the relative positions of the key skeletal points in the image coordinate system form posture change features. Contact perception data is used to form skin pressure distribution and estimate contact stiffness. Skin pressure distribution is used to extract pressure peaks, pressure gradients, and contact areas, while contact stiffness is used to characterize the displacement response after force is applied. Internal state data is used to form fluctuation features, which include at least heart rate fluctuations and blood oxygen saturation fluctuations. Heart rate fluctuations are represented by relative change amplitudes. When heart rate fluctuations exceed 10% or blood oxygen saturation is below 95%, the interaction object is judged to be in an uncomfortable state, and the current state information is recorded.

[0051] In the interactive object state evaluation stage, posture features, pressure distribution features, and fluctuation features are input into the support vector machine classifier to output the limb activity ability level, which is represented by levels from zero to five. At the same time, based on the pressure distribution features and contact stiffness, the skin vulnerability level is output and a skin tolerance threshold is given. The skin tolerance threshold is represented by Newtons and can fall within the range of zero to thirty Newtons. The limb activity angle and the skin tolerance threshold together constitute the mechanical safety boundary input item in the current state information. The current state information also includes an discomfort state marker for adjusting the upper limit of the action force during subsequent constraint synthesis.

[0052] In the semantic recognition stage of nursing objects, a fused feature sequence is constructed based on visual perception data and tactile perception data. This sequence is then input into a combined model of a convolutional network and a long short-term memory network to output the nursing object type and a set of object attributes. The set of object attributes includes at least weight, material, and fragility. Weight is calculated using normal pressure and attitude angle, and the calculation expression adopts...

[0053] Where m is the weight of the patient being cared for. The normal pressure collected by the contact sensing unit. It is the acceleration due to gravity. The angle between the robot's end effector and the horizontal plane is calculated from the robot's end effector posture. The weight calculation result and visual contour features are used together to verify the type of care object, and the grasping force threshold and placement accuracy requirements are written into the care object constraint item in the constraint conditions.

[0054] In the semantic analysis of the nursing environment, depth information from visual perception data is used to generate a 3D point cloud. This 3D point cloud is then processed by random sampling consistency fitting to segment planar and cylindrical structures and label the spatial location and category of medical equipment. Simultaneously, the 2D visual perception data is fed into a semantic recognition network to output environmental semantic labels. These labels include at least infusion status, nursing area, and narrow passage. Based on the equipment location and category, equipment avoidance distances are calculated to form equipment avoidance path constraints. Based on the passable space, the feasible width of the passage is calculated to form action space constraints. All of this information is written into the environmental constraint terms of the constraint conditions.

[0055] In the constraint integration phase, the current state information of the interactive object, the set of nursing object attributes, and the environmental semantic tags are jointly mapped to a constraint set C, which is constrained by the individual patient. Constraints on the care recipients Environmental constraints Composition. Patient individual constraints. Includes the upper limit of the force of movement and the range of limb movement angles, and constraints on the patient being cared for. Includes gripping force threshold and placement accuracy requirements, environmental constraints. It includes equipment avoidance paths and motion space constraints. The constraint set C is written into the constraints of the working environment, so that the constraints of the working environment have the ability to express three types of constraints: individual differences, object attributes and environmental semantics, and maintain a homologous relationship with the feature association information for easy subsequent calling.

[0056] This embodiment extracts skeleton posture features from visual perception data, pressure distribution features from contact perception data, and fluctuation features from internal state data. It then performs semantic-level fusion on these three types of features to generate feature association information. Based on the feature association information, it parses the current state information and generates constraints. This enables the association modeling and interpretable output of multi-source information within the same semantic framework. This transforms the state judgment of interactive objects and the construction of environmental constraints from being driven by scattered data to being driven by association information. It reduces the risk of state misjudgment and constraint loss caused by modal fragmentation and improves the adaptability of subsequent action generation to differences in interactive objects and environmental changes.

[0057] In one embodiment, step S30 above includes: S301, parse the current state information of the interactive object to determine the mechanical safety boundary, and transform the constraints of the working environment into geometric space boundaries; S302, guided by the preset task objectives, establishes a multi-objective cost function that includes safety indicators, comfort indicators, and efficiency indicators; S303, Based on the mechanical safety boundary, the geometric space boundary, and the multi-objective cost function, an optimization model is constructed; S304, The optimization model is solved using a heuristic optimization method to obtain the optimal joint spatial motion sequence and end contact force threshold; S305, the joint space motion sequence is processed by spline curve fitting to generate a continuous motion path, and the end contact force threshold is converted into a control parameter. The continuous motion path and the control parameter are combined to generate an action strategy.

[0058] In this embodiment, the current state information serves as the state input for optimization. It contains at least fields reflecting the physical resilience and state fluctuation of the interactive object. Therefore, the parsing action maps the state semantics into a set of mechanical constraints that can be used to constrain the solution space. The mechanical safety boundary defines the permissible contact force range, force change rate range, and attitude-related force restriction intervals during the interaction process. The determination process maps the physical vulnerability level to tiered force upper limits and force gradient upper limits, and real-time stability to more conservative force change rate thresholds and disturbance tolerance thresholds, thus forming a mechanical constraint expression that changes with the state of the interactive object. The constraints of the operating environment define the feasible domain of the execution space. The process of converting these constraints into geometric space boundaries requires converting semantic descriptions such as obstacle distribution and traversable space into geometric constraints. Geometric space boundaries can be expressed using polyhedral inequalities, occupied grid boundaries, distance field threshold boundaries, or channel boundary curves, enabling subsequent solutions to geometrically exclude inaccessible and interference regions while retaining a safe distance margin.

[0059] The task objective defines the target state and constraint preferences that the planning result needs to achieve, and its role is reflected in the construction of the multi-objective cost function. The multi-objective cost function unifies safety, comfort, and efficiency indicators into scalar or vector objectives that can be processed by the optimization model. Safety indicators measure the degree to which the action process violates mechanical safety boundaries and geometric space boundaries, and can be composed of penalties such as contact force exceeding limits, contact force change rate, minimum distance to obstacles, and path crossing prohibited areas. Comfort indicators measure the smoothness of the action process and the degree of posture disturbance to the interacting object, and can be composed of penalties such as joint acceleration, end-effector velocity mutation, posture change amplitude, and contact force fluctuation, making the solution tend to generate smooth and force-stable actions. Efficiency indicators measure the task completion speed and resource consumption level, and can be composed of penalties such as path length, execution time, energy consumption approximation, or joint motion amplitude, making the solution reduce redundant motion while satisfying safety and comfort. The establishment of the cost function requires unifying the dimensions and scales of the three types of indicators, which can be achieved through normalization, weight coefficient configuration, or hierarchical penalty coefficient configuration, so that the trade-off between different objectives is controllable and adjustable.

[0060] The optimization model couples the mechanical safety boundary, geometric space boundary, and multi-objective cost function into a single solution framework. The model structure includes at least decision variables, a set of constraints, and an objective function. Decision variables can be selected as joint space motion sequences and force control variables related to end-effector contact, enabling the model to simultaneously determine robot joint motion and contact force control. The constraint set consists of the mechanical safety boundary and the geometric space boundary. The mechanical safety boundary constraints restrict the contact force variables and their variation trends, while the geometric space boundary constraints prevent the spatial positions corresponding to the end-effector trajectory and joint trajectory from entering forbidden regions and satisfying safe distances. The objective function is given by the multi-objective cost function, and the model achieves comprehensive optimization of safety, comfort, and efficiency by minimizing the objective function. Since there is a kinematic mapping relationship between joint space motion and spatial geometric constraints, the optimization model needs to include a positive kinematic mapping or an equivalent pose constraint mapping during construction, ensuring that the geometric space boundary can act on the end-effector trajectory corresponding to the joint space motion sequence.

[0061] Heuristic optimization methods are used to search for feasible and cost-effective solutions under complex constraints and multi-objective conditions. The solution output includes the optimal joint spatial motion sequence and end-effector contact force threshold. The joint spatial motion sequence provides a sequence of joint target values ​​or joint increments on discrete time slices. Its discrete granularity and time step can be aligned with the control cycle, enabling the direct generation of executable instructions. The end-effector contact force threshold provides the upper limit or target interval boundary for contact force control. It can be derived from the upper limit value in the mechanical safety boundary, modified by the task objective and comfort constraints, ensuring that it does not exceed the tolerance of the interactive object while meeting the task objective's requirements for contact stability. Heuristic optimization methods can employ sampling-based path search, population-based parameter search, or iterative improvement based on local neighborhoods to achieve the path. In each iteration, feasibility checks and cost assessments are performed. Collision and distance constraint filtering is performed using geometric spatial boundaries, and force feasible domain filtering is performed using mechanical safety boundaries, thereby gradually approaching a better combination of motion sequence and force threshold.

[0062] Spline curve fitting is used to convert discrete joint spatial motion sequences into continuous, executable motion paths. The fitting process constructs piecewise splines using joint angle sequences as control point sequences, and continuity constraints ensure the continuity of position, velocity, and acceleration between segments, resulting in a continuous motion path. This continuous motion path characterizes the continuous trajectory function of the end effector or joint in the time domain, allowing the actuator to obtain target values ​​at any given time within the control cycle through interpolation, reducing the impact of discrete jumps. The process of converting the end effector contact force threshold into control parameters forms a parameter set directly usable by the controller. These control parameters can include upper force limits, force target range parameters, force change rate parameters, or piecewise force parameters associated with the motion path, enabling contact control to employ different force constraint intensities at different path stages. The motion strategy is formed by combining the continuous motion path and control parameters. This combination reflects the coordination of path targets and force targets on the same time axis. The motion strategy can be expressed as a binding structure between a time-parameterized trajectory and a parameter table, allowing the execution phase to simultaneously obtain motion and force control references.

[0063] For example, the system receives nursing task instructions and performs nursing task parsing. These instructions characterize nursing task types such as transfer, turning over, and delivery, as well as the target location or object identifier. The parsing outputs the task type T, target location P, action completion time limit, and safety level S. The safety level S defines the weight range and constraint strength of the safety index in the cost function. Based on the task type T and safety level S, a suitable set of nursing action primitives is selected from the action primitive library. This library includes gripping, lifting, turning over, and grasping primitives. These nursing action primitives define the structural template for joint space motion sequences and carry a set of core parameters. The core parameter set includes at least the action angle θ, execution speed v, contact pressure F, and path trajectory L. The path trajectory L defines the desired motion pattern and key control point distribution of the end effector in three-dimensional space.

[0064] In the action parameter optimization stage, a multi-objective optimization function J is constructed, and the optimal values ​​of the core parameter set are searched using particle swarm optimization. The optimization function is expressed as follows:

[0065] in, , , The weighting coefficients satisfy the following conditions: + + =1, the weighting coefficient is used to make a weighted trade-off between comfort, safety and efficiency, and can increase as the safety level S increases. ; The peak or predicted peak value of contact pressure obtained from tactile information collected during the execution of the action; The contact pressure threshold that the interactive object can withstand; The minimum safe distance or predicted minimum safe distance between the trajectory of the action and the obstacle; This serves as the upper bound for the normalization of safe distances, used to standardize the scale of safe distances in different scenarios. Estimate the execution time for the action; This serves as the task completion time limit or upper time bound. During particle swarm optimization, particle positions are encoded as control point parameters θ, v, F, and path trajectory L. Based on the minimization result of J, the optimized motion angle, execution speed, contact pressure, and path trajectory parameters are output, and the joint spatial motion sequence and end-effector contact force threshold are determined accordingly.

[0066] In the trajectory generation stage, a continuous motion path is generated using B-spline curve fitting based on the optimized path trajectory parameters. This ensures continuous velocity and acceleration and limits abrupt curvature changes to reduce the risk of impact during nursing care. For nursing tasks requiring precise positioning, pause points are inserted into the motion path with set dwell times, such as a 0.5s pause near the target area to confirm location and verify object status. Finally, the end-effector contact force threshold is converted into control parameters and combined with the continuous motion path to obtain the motion strategy output. The control parameters constrain the force control target or amplitude limit of the motor drive command, while the motion path constrains the trajectory following target and timing execution rhythm of the motor drive command.

[0067] This embodiment parses the current state information into a mechanical safety boundary and transforms the constraints into a geometric space boundary. Then, guided by the task objective, it establishes a multi-objective cost function that includes safety, comfort, and efficiency indicators, and constructs an optimization model accordingly. Combining heuristic optimization methods, it outputs the joint spatial motion sequence and end-effector contact force threshold. Then, it generates a continuous motion path through spline curve fitting and transforms the end-effector contact force threshold into control parameters to form a motion strategy. This achieves a unified expression of state constraints, spatial constraints, and multi-objective trade-offs within the same solution framework, enabling the motion strategy to simultaneously possess path reachability and force controllability under feasible domain constraints. This reduces the risk of infeasible planning and force exceeding limits caused by state differences and environmental limitations, and improves trajectory continuity and the executability of control parameters.

[0068] In one embodiment, step S40 above includes: S401, the control parameters and motion path contained in the motion strategy are converted into motor drive commands, and the motor drive commands are sent to the underlying controller of the robot execution device to drive the robot execution device to move; S402 uses a force sensor installed at the operating end of the robot actuator to collect the actual contact pressure value in real time; S403, Read the joint encoder data of the robot actuator and determine the actual pose coordinates of the operating end of the robot actuator based on the joint encoder data; S404, record the real-time intrinsic state fluctuation value of the interactive object, and summarize the actual contact pressure value, the actual pose coordinates and the real-time intrinsic state fluctuation value into execution feedback data.

[0069] In this embodiment, the motion strategy provides control content that can be directly issued to the drive link during the execution phase. The control parameters bear the target constraints and amplitude limits related to force control, while the motion path bears the reference trajectory of the pose changing over time. The process of converting the control parameters and motion path into motor drive commands requires two types of mapping: one is the mapping from trajectory to joint space, where the motion path is sampled into an end-effector pose sequence over discrete control cycles based on the robot's kinematic model, and then the joint target sequence is obtained through inverse kinematics or Jacobian iteration, further generating position or velocity mode drive quantities; the other is the mapping from control parameters to drive amplitude limits, converting constraints such as end-effector contact force thresholds and force change rate limits into torque limits, velocity limits, acceleration limits, or impedance parameter tables, ensuring that the underlying controller satisfies force constraints while executing the joint target sequence. Structurally, the motor drive command can consist of joint target values, joint target velocities, joint target torques, amplitude limits, and control cycle identifiers, enabling the underlying controller to perform interpolation, closed-loop adjustment, and safety limiting within each control cycle.

[0070] Motor drive commands are sent to the underlying controller to form a closed-loop interface from planning to execution. The transmission process requires time alignment information and a sequence index to ensure the underlying controller consumes commands according to the predetermined control cycle and maintains trajectory continuity. To reduce the impact of communication jitter on execution, the transmitting side can employ a ring buffer and advance delivery mechanism. This ensures the underlying controller always has at least one interpolation window length of available commands, and in the event of a missing command, it reverts to the most recent valid command while maintaining the amplitude constraint.

[0071] Force sensors are used to acquire actual contact pressure values ​​in real time. The acquisition link needs to complete sensor zero-point calibration, range conversion, and sampling synchronization. The definition of the actual contact pressure value needs to establish a definite correspondence with the sensor output. When the sensor output is a multi-axis force or array pressure, the pressure scalar can be obtained by extracting the normal component, integrating the region, or aggregating the maximum value, while retaining the sampling timestamp to support alignment with the sampling points along the motion path. Real-time performance requires that the sampling period not exceed the underlying control period so that subsequent deviation calculations can reflect rapid changes in the contact state.

[0072] Joint encoder data is used to characterize the actual position state of each joint of the robot's actuator. The reading process needs to be aligned with the underlying control cycle and includes filtering and outlier removal to avoid encoder jitter causing abrupt changes in pose calculation. Determining the actual pose coordinates of the end effector based on the joint encoder data requires performing forward kinematics calculations. The joint angles are substituted into the link parameter model to output the end effector pose. Simultaneously, calibration compensation parameters are introduced as needed to correct zero-position deviations and link parameter errors. The actual pose coordinates need to be consistent with the reference coordinate system of the motion path. Therefore, a coordinate transformation matrix can be added after calculation to unify the results to the reference coordinate system of the execution stage, ensuring that pose comparisons have the same geometric reference.

[0073] Real-time intrinsic state fluctuation values ​​describe the state fluctuations of the interactive object during execution. The source can be a continuously pushed numerical sequence or event sequence from the data interface. The recording process needs to be aligned with the control cycle and include rules for handling missing values, ensuring that the intrinsic state fluctuation values ​​can be concatenated with the actual contact pressure value and actual pose coordinates on the timeline. Summarizing the actual contact pressure value, actual pose coordinates, and real-time intrinsic state fluctuation values ​​into execution feedback data requires a unified data structure, including at least a timestamp, pressure field, pose field, state fluctuation field, and quality marker field. The quality marker field records whether the sampling is valid, whether it exceeds the range, and whether packet loss occurs, thus providing a verifiable data foundation for subsequent deviation calculations and safety monitoring.

[0074] For example, the execution drive module receives control parameters and motion paths output by the motion strategy, encodes the target angle, target speed, target contact pressure, and discrete trajectory point sequence of the motion path into motor drive commands, and sends them to the underlying controller to drive the joint motor and end effector to move in coordination. In a patient transfer scenario, the underlying controller controls the pose change and movement speed of the support arm according to the motion path, and simultaneously controls the support force and fine-tuning of the support position according to the target contact pressure. The tactile and visual fusion perception module continuously acquires execution feedback data during motion execution. The tactile side aggregates pressure changes in the end-effector contact area, while the visual side extracts limb posture deviation and end-effector offset relative to the target position, and detects path restriction or changes in passable space caused by dynamic environmental interference. In the feedback analysis phase, the deviation between the actual motion parameters and the planned parameters is calculated based on the execution feedback data, where the pressure deviation can be expressed as... The angular deviation can be expressed as F represents the actual contact pressure value in the execution feedback data, in Newtons. The target contact pressure in the control parameters is expressed in Newtons, and θ is the actual angle corresponding to the execution feedback data. To control the target angle value in the parameters, the angle value can be selected as joint angle or end-effector attitude angle, in degrees. and These represent haptic control deviation and posture control deviation, respectively. Based on the deviation results, the control output adjustment module uses fuzzy inference logic in conjunction with the PID control loop to generate real-time correction values ​​and update the motor drive commands.

[0075] When | When the pressure exceeds a preset tolerance, such as 1N, reduce the target contact pressure of the end effector or reduce the force control gain to decrease the output force. | When the angle exceeds the preset tolerance, for example, 2°, incremental correction is performed on the joint target angle or end-effector posture target to reduce posture deviation. When the visual side detects micro-movement of the interactive object's limb, or when the heart rate corresponding to the real-time intrinsic state fluctuation value in the execution feedback data suddenly increases and reaches the preset judgment condition, the underlying controller enters the pause control mode and maintains the current position. At the same time, the current deviation and environmental state snapshot are sent back to the optimization solution stage to trigger the recalculation of control parameters. For high-risk nursing actions such as wound care and fracture transfer, action pre-playing is added before execution. The accessibility of the motion path, joint limits, upper limit of end-effector contact pressure, and potential collision risks are verified using a virtual simulation environment. Only after the verification is passed will the motor drive command issuance and closed-loop adjustment process begin. If the verification fails, the parameter combination of target angle, target speed, or target contact pressure is adjusted and the motor drive command is regenerated.

[0076] This embodiment converts control parameters and motion paths into motor drive commands and sends them to the underlying controller, realizing an executable mapping from motion strategy to execution link. It collects actual contact pressure values ​​through force sensors, determines actual pose coordinates through joint encoder data, and simultaneously records real-time internal state fluctuation values. These three are then summarized into execution feedback data along the time axis, enabling real-time observation of force, pose, and interactive object state fluctuations during the execution phase. This improves the completeness and time consistency of execution feedback data, providing continuous, aligned, and traceable data support for subsequent deviation calculation, control output adjustment, and safety monitoring.

[0077] In one embodiment, step S50 above includes: S501, the theoretical force control value corresponding to the control parameters and the theoretical motion trajectory corresponding to the motion path are parsed from the action strategy; S502, determine the pressure deviation between the actual contact pressure value and the theoretical force control value contained in the execution feedback data, and determine the trajectory deviation between the actual pose coordinates and the theoretical motion trajectory contained in the execution feedback data; S503, taking the pressure deviation and the trajectory deviation as inputs, the torque correction amount and the angle correction amount are determined by fuzzy reasoning logic; S504, the torque correction amount and the angle correction amount are superimposed on the current motor drive command in real time to dynamically adjust the control output of the robot execution device.

[0078] In this embodiment, the execution feedback data is used to characterize the actual force and pose states of the robot's execution device during operation, while the motion strategy is used to characterize the target force constraints and target motion trajectory given in the planning phase. To determine the execution deviation between the execution feedback data and the control parameters and motion paths in the motion strategy, the control parameters and motion paths in the motion strategy need to be converted into comparable reference quantities. Furthermore, the actual contact pressure values ​​and actual pose coordinates in the execution feedback data should be aligned to these reference quantities along the same time axis and coordinate system to ensure that the deviation calculation has the same reference benchmark.

[0079] To parse the theoretical force control values ​​corresponding to the control parameters from the motion strategy, it is necessary to establish a mapping relationship between the control parameters and the force target. Control parameters may include fields such as end-contact force threshold, force change rate limit, force control direction constraint, and impedance parameters. The theoretical force control value can be directly determined from the end-contact force threshold, or the equivalent force target can be calculated from the relationship between the impedance parameters and the desired displacement. Furthermore, the theoretical force control value needs to be bound to the sampling points of the motion path or the control cycle index, so that the theoretical force control value and the actual contact pressure value can be compared within the same control cycle. The parsing process can output a sequence of theoretical force control values ​​indexed by the control cycle, thus providing a cycle-by-cycle reference for subsequent pressure deviation calculations.

[0080] To deduce the theoretical motion trajectory from the motion strategy, the reference curve representing the motion path needs to be discretized into a directly comparable pose sequence or pose increment sequence. The motion path can be an end-effector pose curve, a keypoint sequence, or a parametric spline curve. The theoretical motion trajectory can be obtained by sampling the target pose sequence at equal intervals according to the control cycle, while maintaining a spatial coordinate system definition and attitude representation consistent with the actual pose coordinates. The theoretical motion trajectory can also include velocity or acceleration references to support multi-dimensional measurements of trajectory deviations, such as joint measurements of position and attitude deviations, or combined measurements of position, attitude, and velocity deviations.

[0081] Determining the pressure deviation between the actual contact pressure value and the theoretical force control value requires defining the calculation rules and tolerance rules for the pressure deviation. The pressure deviation can be expressed as the difference between the actual contact pressure value and the theoretical force control value, or it can be expressed as a normalized form to adapt to different ranges and contact stages. To ensure the stability of the pressure deviation, a sliding window aggregation of the actual contact pressure values ​​can be performed within the same control cycle to obtain a representative value for deviation calculation. The deviation sign is retained to distinguish between excessive and insufficient pressure, thus enabling subsequent control output adjustments to differentiate between the direction of force reduction and the direction of force increase. The time index of the pressure deviation needs to be consistent with the theoretical force control value sequence to avoid spurious deviations introduced by sampling delays.

[0082] Determining the trajectory deviation between the actual pose coordinates and the theoretical motion trajectory requires matching the corresponding sampling points of the actual pose coordinates and the theoretical motion trajectory, and then calculating the geometric error after matching. Matching methods can be based on a one-to-one correspondence using the control cycle index, or on nearest-point matching based on the arc length parameter, thus adapting to phase shifts caused by velocity changes. The trajectory deviation can include two components: position deviation and attitude deviation. Position deviation can be expressed using three-dimensional vector difference or Euclidean distance, while attitude deviation can be expressed using quaternion distance, rotation vector, or Euler angle difference. Furthermore, the trajectory deviation needs to be consistent with the control output form of the robot's actuator so that subsequent corrections can directly affect the motor drive commands. The trajectory deviation can also be represented in joint space by mapping the theoretical motion trajectory and the actual pose coordinates to joint angle sequences, and then calculating the joint angle deviation to support the generation of angle corrections.

[0083] Using pressure deviation and trajectory deviation as inputs and determining torque and angle corrections through fuzzy inference logic, continuous deviations need to be mapped to discrete semantic sets, and a rule set needs to be established. The fuzzy inference logic can construct membership functions for pressure and trajectory deviations respectively. These membership functions map deviation amplitudes to membership degrees of multiple fuzzy sets; for example, multi-level deviation semantic sets can be formed based on amplitude intervals, and membership degrees can be defined for the positive and negative directions of pressure deviation to distinguish between force increase and decrease control directions. The rule set describes the correspondence between the combined states of pressure and trajectory deviations and control corrections. The rule output can contain fuzzy sets of torque and angle corrections. After aggregation and defuzzification, continuous numerical torque and angle corrections are obtained. Torque corrections can be expressed as the torque increment of each joint or the equivalent torque increment at the end effector, and angle corrections can be expressed as the angle increment of each joint or the attitude correction increment at the end effector. Both need to maintain the same dimensions and units as the motor drive command fields received by the underlying controller.

[0084] The torque and angle correction values ​​are superimposed on the current motor drive command in real time to form a closed-loop adjustment during execution. The superposition action requires a clear superposition position and order. The torque correction value can be superimposed on the torque target field or torque limit field of the motor drive command, and the angle correction value can be superimposed on the joint target position field or joint target increment field of the motor drive command. Furthermore, the superposition process must satisfy control cycle constraints to maintain continuity. To avoid abrupt changes in control commands due to superposition, amplitude and rate of change limits can be imposed on the torque and angle correction values, ensuring that the correction increment in each control cycle does not exceed a preset upper limit. This maintains the trackability of the motion path and the controllability of the force state while dynamically adjusting the control output.

[0085] This embodiment obtains theoretical force control values ​​and theoretical motion trajectories from motion strategy analysis, and aligns them with the actual contact pressure values ​​and actual pose coordinates in the execution feedback data to form quantifiable deviation representations of pressure deviation and trajectory deviation. Furthermore, using pressure deviation and trajectory deviation as inputs, fuzzy inference logic is used to generate torque correction and angle correction values, which are then superimposed on the motor drive commands in real time to achieve closed-loop dynamic adjustment of the control output. This enables the robot actuator to obtain continuous deviation suppression capability when both the force target and trajectory target exist simultaneously, reducing the cumulative risk of force overshoot and trajectory drift.

[0086] In one embodiment, step S60 above includes: S601, based on the physical limitations included in the constraints of the working environment, a first-level pressure safety threshold, a second-level pressure safety threshold, and a third-level pressure safety threshold are set for contact pressure, and a first-level trajectory safety threshold, a second-level trajectory safety threshold, and a third-level trajectory safety threshold are set for trajectory deviation. The first-level pressure safety threshold and the first-level trajectory safety threshold correspond to parameter fine-tuning operations, the second-level pressure safety threshold and the second-level trajectory safety threshold correspond to action deceleration operations, and the third-level pressure safety threshold and the third-level trajectory safety threshold correspond to emergency braking operations. S602, real-time acquisition of the actual contact pressure value and actual pose coordinates in the execution feedback data, determination of the trajectory deviation between the actual pose coordinates and the motion path in the action strategy, determination of whether the actual contact pressure value exceeds the first level pressure safety threshold, the second level pressure safety threshold or the third level pressure safety threshold, and determination of whether the trajectory deviation exceeds the first level trajectory safety threshold, the second level trajectory safety threshold or the third level trajectory safety threshold; S603, if the actual contact pressure value or the trajectory deviation exceeds the threshold of the corresponding level, then perform the corresponding parameter fine-tuning operation, action deceleration operation or emergency braking operation according to the threshold level exceeded. S604, after the action strategy is executed, the frequency of pressure exceeding the limit, trajectory smoothness and internal state recovery time in the execution feedback data are statistically analyzed to generate effect analysis information; S605, the effect analysis information is used as a reward signal for reinforcement learning, and the model parameters of the optimization model are adjusted using the reinforcement learning module.

[0087] In this embodiment, safety monitoring and response are based on two types of inputs: constraints and execution feedback data. Constraints provide an acceptable range of physical limitations for the operating environment. These limitations can cover upper limits on contact pressure, allowable trajectory deviations, movement restrictions within specific areas, and force-restricted zones related to the interacting object, thus providing an external constraint source for threshold setting. Execution feedback data characterizes the actual force and pose state of the robot's actuator during operation, including at least the actual contact pressure value and actual pose coordinates, thereby providing direct observations for pressure-related and trajectory-related determinations.

[0088] The tiered safety thresholds are set based on contact pressure and trajectory deviation, respectively. The first, second, and third level pressure safety thresholds together constitute the pressure tiered intervals. The boundaries of these intervals are derived from the dimensionality standardization of physical limits. This standardization can include unit conversion, sensor calibration coefficient conversion, and equivalent pressure calculation after normalizing the end-contact area, allowing direct comparison of the thresholds with actual contact pressure values. The first, second, and third level trajectory safety thresholds together constitute the trajectory deviation tiered intervals. The boundaries of these intervals are derived from geometric space limitations and allowable offset ranges. The allowable offset range can be determined based on the boundaries of the passageway, obstacle safety distances, and the permissible swing range of the end-contact area, allowing direct comparison of trajectory deviations with the trajectory safety thresholds. The pressure and trajectory thresholds are mapped to parameter fine-tuning operations, deceleration operations, and emergency braking operations, respectively. This mapping relationship converts numerical judgment results into executable control action levels, avoiding binary decisions of simply outputting "stop" or "continue."

[0089] To acquire the actual contact pressure value and actual pose coordinates in real time, it is necessary to control the sampling frequency and timestamp of the execution feedback data consistently, ensuring that pressure sampling and pose sampling fall within the same control cycle. The actual contact pressure value can be obtained from the force sensor output through filtering and zero-point drift compensation, ensuring that the pressure value used for threshold comparison is not dominated by transient noise. The actual pose coordinates can be calculated by the joint encoder or external positioning and mapped to a spatial coordinate system consistent with the motion path, ensuring that the same reference system is used for subsequent trajectory deviation calculations.

[0090] The determination of trajectory deviation is based on the correspondence between the actual pose coordinates and the motion path in the motion strategy. The motion path is typically expressed as a parametric curve or discrete path points. Trajectory deviation requires converting the motion path into a theoretical pose sequence and matching it with the actual pose coordinates. Matching methods can include alignment based on the control cycle index or nearest-point matching based on the path arc length to handle phase differences caused by velocity variations. Trajectory deviation can be measured using a combination of position and attitude deviations. Position deviation can be expressed using three-dimensional Euclidean distance or partial axis deviation, while attitude deviation can be expressed using rotation vector difference or quaternion distance, and can be further converted into a single scalar deviation for direct comparison with the trajectory safety threshold. After calculation, the trajectory deviation enters the threshold determination link, forming the safety trigger condition for the trajectory direction.

[0091] Threshold determination needs to cover both pressure direction and trajectory direction. The determination of actual contact pressure value can employ a hierarchical comparison logic to identify the pressure's range or the highest threshold it exceeds, thus generating a pressure trigger level. Similarly, the determination of trajectory deviation uses a hierarchical comparison logic to identify the deviation's range or the highest threshold it exceeds, thus generating a trajectory trigger level. These two determination chains operate in parallel; triggering either chain can initiate a safety response. To avoid action conflicts when pressure and trajectory trigger simultaneously, a level synthesis rule can be used. For example, the higher of the pressure trigger level and the trajectory trigger level can be taken as the final response level, giving priority to emergency braking, intermediate-level response to deceleration, and minimum-level response to parameter fine-tuning.

[0092] The execution of a safety response is based on the response level as input. Parameter fine-tuning focuses on correcting control parameters without altering the mission trajectory. This can include lowering the contact force target, adjusting impedance parameters, adjusting the end-effector velocity limit, and injecting attitude fine-tuning increments to bring pressure or trajectory deviations back within thresholds. Action deceleration focuses on reducing system kinetic energy and the rate of deviation growth. This can include speed scaling, acceleration limiting tightening, and path-following gain reduction to enable a more robust following state. Emergency braking focuses on terminating increased risk. This can include issuing a stop command, setting force control output to zero or switching to a safety hold mode, and maintaining the end-effector attitude without further propulsion. The trigger condition for emergency braking corresponds to the third-level pressure safety threshold or the third-level trajectory safety threshold, ensuring a clear response action even in the highest-risk scenarios.

[0093] The generation of performance analysis information occurs after the action policy is executed, with the goal of converting execution feedback data into evaluation metrics that can be used for model updates. Stress over-limit frequency characterizes the number of times the stress safety threshold is triggered or the number of over-limit events. Statistics can be based on edge detection of threshold over-limit events to avoid double counting during continuous over-limit periods. Trajectory smoothness characterizes the continuity and jitter level of the actual pose trajectory. Trajectory smoothness can be calculated from indicators such as the second-order difference of the pose sequence, the rate of change of curvature, and the rate of change of velocity direction, reflecting the quality of trajectory execution. Intrinsic state recovery time characterizes the speed at which the interactive object's state falls back to normal. Intrinsic state fluctuation values ​​can be derived from state observations related to the interactive object in the execution feedback data, and recovery time can be defined as the length of time required for the fluctuation value to fall back to the target interval. The above statistical results are combined to form performance analysis information. Performance analysis information needs to maintain a structured expression, including at least three types of fields: frequency, smoothness, and recovery time, so that it can be subsequently used as a reward signal input to the reinforcement learning module.

[0094] The reinforcement learning module uses performance analysis information as a reward signal. This requires mapping the performance analysis information into numerical rewards and establishing a correlation between this information and the update of the optimized model parameters. The reward signal can be obtained by weighting multiple indicators. The frequency of stress exceeding limits can be used as a penalty, trajectory smoothness as a positive or negative indicator, and intrinsic state recovery time as a penalty. Weights can be set according to the safety priority principle, ensuring the reward signal reflects constraint priority. Adjustments to model parameters must be limited to the adjustable parameter set of the optimized model. This set can include cost function weights, constraint relaxation coefficients, mapping coefficients for control parameter generation, and smoothing term weights for path generation, allowing the update results to alter the generation preferences of action policies in subsequent solution stages. The update process can employ batch updates or online updates. Batch updates aggregate rewards based on data from multiple execution rounds, while online updates update parameters based on a single execution round. Both methods use reward signals to drive model parameter adjustments along the direction of increasing rewards, thus forming a closed loop of safety monitoring, response, evaluation, and update.

[0095] For example, during the execution of nursing actions, the safety closed-loop monitoring module receives multi-source feedback data from touch, vision, and physiology, forming a unified monitoring data stream. On the tactile side, the force sensor outputs the actual contact pressure value; on the visual side, pose calculation outputs the actual pose coordinates and further obtains the trajectory deviation; on the physiological side, vital sign monitoring outputs real-time intrinsic state fluctuation values, which are mapped to the fluctuation amplitude of indicators such as heart rate, blood pressure, and blood oxygen saturation. The physical limitations included in the constraints are interpreted as upper limits for contact pressure and trajectory deviation, and a graded safety threshold set is generated accordingly. The upper limit for contact pressure is denoted as... The upper limit of trajectory deviation is denoted as Level 1 pressure safety threshold =0.8 Second-level pressure safety threshold =0.9 Level 3 pressure safety threshold = Level 1 trajectory safety threshold =0.8 Second-level trajectory safety threshold =0.9 Level 3 trajectory safety threshold = ,in This indicates the maximum permissible contact pressure, expressed in Newtons. This indicates the maximum allowable trajectory deviation; the unit can be selected as meters or millimeters. to and to These represent the grading thresholds for contact pressure and trajectory deviation, respectively.

[0096] During the real-time monitoring phase, the safety closed-loop monitoring module reads the actual contact pressure value and calculates the trajectory deviation at a fixed refresh cycle. The trajectory deviation is calculated by comparing the actual pose coordinates with the motion path in the action strategy. The motion path can be discretized into a reference pose sequence, and the distance from the current position to the reference pose is taken as the deviation amplitude. The monitoring module then determines whether the actual contact pressure value exceeds the limit. , , And whether the trajectory deviation exceeds , , The system maintains a mapping relationship between thresholds and response actions within its tiered rules. Level 1 thresholds trigger parameter fine-tuning, level 2 thresholds trigger deceleration actions, and level 3 thresholds trigger emergency braking actions. When physiological indicators in the monitored data stream need to participate in early warning, the heart rate fluctuation amplitude is defined as... ,in This is the current heart rate, expressed in beats per minute. The baseline heart rate value is the average of the resting heart rate or the heart rate within a stable window before initiating nursing intervention. As the relative fluctuation ratio, when When the preset ratio threshold is reached, such as 0.15, the event is associated with a level 2 or level 3 response action to reduce the risk of impact.

[0097] During the emergency risk handling phase, in addition to threshold triggering, the monitoring module also establishes event judgment conditions for faults and external disturbances. Robot actuator faults are identified through fault codes, abnormal drive currents, or encoder sync failures in the underlying controller, triggering the switching of redundant control channels and rerouting motor drive commands to backup actuators or backup drive channels to maintain the support posture. Interaction object anomalies are jointly judged by real-time internal state fluctuation values ​​and posture changes on the vision side, triggering support posture adjustments and sending alarm messages to the alarm terminal to prompt nursing staff to intervene. When obstacles appear in the environment, the changes in obstacle distribution on the vision side are judged and a motion path regeneration request is triggered. The current pose, obstacle distribution, passable space, and constraints are encapsulated as replanning inputs to avoid path interference. When emergency braking is triggered, audible and visual alarms or network alarms are activated simultaneously to shorten the manual response link.

[0098] In the effectiveness evaluation phase, after the action strategy is executed, multidimensional statistics are extracted from the execution feedback data to form effectiveness analysis information. Safety indicators include the frequency of stress exceeding limits and the number of action interventions. The frequency of stress exceeding limits can be calculated as follows: Statistically, F is the actual contact pressure value, and I() is the indicator function. The number of times the heart rate exceeded the limit; comfort indicators include the peak value of heart rate variability. ,max and trajectory deviation peak Efficiency metrics include actual execution time. With planning time Deviation ΔT= ,in The actual time taken to execute the action. The time budget corresponding to the action strategy is used; the effect analysis information and nursing process data are written into the knowledge base to form searchable historical samples. The reinforcement learning module constructs reward signals from the effect analysis information and uses them to adjust and optimize the model parameters, so that the control parameters and motion paths generated in the future under the same or similar constraints and interactive object states are more in line with safety threshold constraints and comfort goals.

[0099] This embodiment constructs pressure grading thresholds and trajectory grading thresholds based on physical constraints, and incorporates actual contact pressure values ​​and trajectory deviations into the grading comparison chain to form a quantifiable and gradable safety triggering mechanism. Furthermore, it maps triggering levels to parameter fine-tuning operations, action deceleration operations, and emergency braking operations, achieving a deterministic correspondence between risk levels and response intensity. After the action strategy is executed, it statistically analyzes the frequency of pressure exceeding limits, trajectory smoothness, and intrinsic state recovery time from the execution feedback data to generate effect analysis information. This effect analysis information is then transformed into reward signals to drive the reinforcement learning module to adjust and optimize the model parameters, enabling subsequent action strategy generation to adaptively correct based on historical safety and execution quality feedback, reducing the probability of continuous accumulation of pressure exceeding limits and trajectory deviations, and improving the stability and consistency of the execution process.

[0100] In one embodiment, a multi-source sensing-based interactive control device is provided, which corresponds one-to-one with the multi-source sensing-based interactive control method described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the interactive control device based on multi-source sensing of the present invention. The modules include a multi-source sensing data fusion module 10, a state and constraint parsing module 20, an action strategy generation module 30, an execution and feedback acquisition module 40, an execution deviation adjustment module 50, and a safety monitoring and model update module 60. Detailed descriptions of each functional module are as follows: The multi-source sensing data fusion module 10 is used to collect multi-source heterogeneous sensing data in the working environment and preprocess the multi-source heterogeneous sensing data to generate a standardized sensing dataset. The state and constraint parsing module 20 is used to extract features and semantically associate the standardized perception dataset to obtain feature association information, and to parse the current state information of the interactive object and the constraints of the working environment based on the feature association information. The action strategy generation module 30 is used to generate an action strategy containing control parameters and motion path by solving the current state information and the constraints, combined with the task objective, using an optimization model. The execution and feedback acquisition module 40 is used to control the robot execution device to run according to the action strategy and to collect execution feedback data in real time during the operation. The execution deviation adjustment module 50 is used to determine the execution deviation between the execution feedback data and the control parameters and motion path in the action strategy, and adjust the control output of the robot execution device according to the execution deviation; The safety monitoring and model update module 60 is used to perform safety monitoring and response based on the constraints and the execution feedback data, and to generate effect analysis information based on the execution feedback data to update the optimization model.

[0101] For specific limitations regarding the multi-source sensing-based interactive control device, please refer to the aforementioned limitations on the multi-source sensing-based interactive control method, which will not be repeated here. Each module in the aforementioned multi-source sensing-based interactive control device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the processor in the computer device, or stored in software in the memory of the computer device, so that the processor can call and execute the operations corresponding to each module.

[0102] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a server-side interactive control method based on multi-source perception.

[0103] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements client-side functions or steps of a multi-source sensing-based interactive control method.

[0104] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Collect multi-source heterogeneous sensing data from the working environment, and preprocess the multi-source heterogeneous sensing data to generate a standardized sensing dataset; Feature extraction and semantic association are performed on the standardized perception dataset to obtain feature association information, and the current state information of the interactive object and the constraints of the working environment are parsed based on the feature association information. Based on the current state information and the constraints, and combined with the task objective, an optimization model is used to solve the problem and generate an action strategy that includes control parameters and motion path. The robot's execution device is controlled to operate according to the aforementioned action strategy, and execution feedback data is collected in real time during the operation. Determine the execution deviation between the execution feedback data and the control parameters and motion path in the action strategy, and adjust the control output of the robot execution device according to the execution deviation; Based on the constraints and the execution feedback data, security monitoring and response are performed, and effect analysis information is generated based on the execution feedback data to update the optimization model.

[0105] In one embodiment, a computer-readable storage medium is provided, which may be non-volatile or volatile, and a computer program is stored thereon, which, when executed by a processor, performs the following steps: Collect multi-source heterogeneous sensing data from the working environment, and preprocess the multi-source heterogeneous sensing data to generate a standardized sensing dataset; Feature extraction and semantic association are performed on the standardized perception dataset to obtain feature association information, and the current state information of the interactive object and the constraints of the working environment are parsed based on the feature association information. Based on the current state information and the constraints, and combined with the task objective, an optimization model is used to solve the problem and generate an action strategy that includes control parameters and motion path. The robot's execution device is controlled to operate according to the aforementioned action strategy, and execution feedback data is collected in real time during the operation. Determine the execution deviation between the execution feedback data and the control parameters and motion path in the action strategy, and adjust the control output of the robot execution device according to the execution deviation; Based on the constraints and the execution feedback data, security monitoring and response are performed, and effect analysis information is generated based on the execution feedback data to update the optimization model.

[0106] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0107] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0108] It should be noted that if any AI models, software tools, or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The above-described embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

[0109] The user personal information involved in this application embodiment is all authorized (knowing and consenting) by the relevant parties or fully authorized by all parties, and the executing entity can obtain it through various open, legal and compliant means. The collection, storage, use, processing, transmission, provision and disclosure of the information, data and signals involved all comply with the relevant laws and regulations of the relevant countries and regions, and do not violate public order and good morals.

Claims

1. An interactive control method based on multi-source sensing, characterized in that, Includes the following steps: Collect multi-source heterogeneous sensing data from the working environment, and preprocess the multi-source heterogeneous sensing data to generate a standardized sensing dataset; Feature extraction and semantic association are performed on the standardized perception dataset to obtain feature association information, and the current state information of the interactive object and the constraints of the working environment are parsed based on the feature association information. Based on the current state information and the constraints, and combined with the task objective, an optimization model is used to solve the problem and generate an action strategy that includes control parameters and motion path. The robot's execution device is controlled to operate according to the aforementioned action strategy, and execution feedback data is collected in real time during the operation. Determine the execution deviation between the execution feedback data and the control parameters and motion path in the action strategy, and adjust the control output of the robot execution device according to the execution deviation; Based on the constraints and the execution feedback data, security monitoring and response are performed, and effect analysis information is generated based on the execution feedback data to update the optimization model.

2. The interactive control method based on multi-source sensing as described in claim 1, characterized in that, Collect multi-source heterogeneous sensing data from the operating environment, and preprocess the multi-source heterogeneous sensing data to generate a standardized sensing dataset, including: The system controls the contact sensing unit located at the end of the robot's execution device to collect contact sensing data, controls the vision acquisition unit to collect visual sensing data of the working environment, and obtains the internal state data of the interactive object through the data interface. Based on the contact sensing data, the visual sensing data, and the internal state data, multi-source heterogeneous sensing data is constructed. Using the hardware trigger signal as a unified time reference, timestamp alignment is performed on the contact perception data, the visual perception data, and the intrinsic state data in the multi-source heterogeneous perception data. Filtering is performed on the timestamp-aligned contact perception data, visual perception data, and internal state data respectively to obtain filtered contact perception data, filtered visual perception data, and filtered internal state data. The filtered contact perception data, the filtered visual perception data, and the filtered intrinsic state data are mapped to a unified spatial coordinate system for data format normalization, generating a standardized perception dataset.

3. The interactive control method based on multi-source perception as described in claim 1, characterized in that, Feature extraction and semantic association are performed on the standardized perception dataset to obtain feature association information. Based on the feature association information, the current state information of the interactive object and the constraints of the working environment are parsed, including: Skeletal pose features are extracted from the visual perception data contained in the standardized perception dataset, pressure distribution features are extracted from the contact perception data contained in the standardized perception dataset, and fluctuation features are extracted from the intrinsic state data contained in the standardized perception dataset. The skeleton posture features, the pressure distribution features, and the wave features are semantically fused to generate feature association information. Based on the feature association information, analyze the physical vulnerability level and real-time stability of the interactive object, and generate the current state information of the interactive object; Based on the feature association information, the distribution of obstacles and passable spaces in the work environment are identified, and based on the obstacle distribution, passable spaces and the current state information of the interactive objects, the constraints of the work environment are generated.

4. The interactive control method based on multi-source perception as described in claim 1, characterized in that, Based on the current state information and the constraints, and in conjunction with the task objective, an optimization model is used to solve the problem, generating an action strategy that includes control parameters and a motion path, including: The current state information of the interactive object is parsed to determine the mechanical safety boundary, and the constraints of the working environment are transformed into geometric space boundaries; Guided by the goal of meeting the preset task objectives, a multi-objective cost function is established, which includes safety indicators, comfort indicators, and efficiency indicators. An optimization model is constructed based on the mechanical safety boundary, the geometric space boundary, and the multi-objective cost function; The optimization model is solved using a heuristic optimization method to obtain the optimal joint spatial motion sequence and end contact force threshold. The joint space motion sequence is processed using spline curve fitting to generate a continuous motion path, and the end contact force threshold is converted into a control parameter. The continuous motion path and the control parameter are combined to generate a motion strategy.

5. The interactive control method based on multi-source perception as described in claim 1, characterized in that, The robot's execution device is controlled to operate according to the aforementioned action strategy, and execution feedback data is collected in real time during operation, including: The control parameters and motion path contained in the motion strategy are converted into motor drive commands, and the motor drive commands are sent to the underlying controller of the robot execution device to drive the robot execution device to move; The actual contact pressure value is collected in real time using a force sensor installed at the operating end of the robot's actuator; Read the joint encoder data of the robot actuator and determine the actual pose coordinates of the robot actuator end effector based on the joint encoder data; Record the real-time intrinsic state fluctuation value of the interactive object, and summarize the actual contact pressure value, the actual pose coordinates and the real-time intrinsic state fluctuation value into execution feedback data.

6. The interactive control method based on multi-source sensing as described in claim 1, characterized in that, Determining the execution deviation between the execution feedback data and the control parameters and motion path in the motion strategy, and adjusting the control output of the robot execution device according to the execution deviation, includes: The theoretical force control values ​​corresponding to the control parameters and the theoretical motion trajectory corresponding to the motion path are extracted from the action strategy. Determine the pressure deviation between the actual contact pressure value and the theoretical force control value contained in the execution feedback data, and determine the trajectory deviation between the actual pose coordinates and the theoretical motion trajectory contained in the execution feedback data; Using the pressure deviation and the trajectory deviation as inputs, the torque correction amount and the angle correction amount are determined through fuzzy reasoning logic. The torque correction amount and the angle correction amount are superimposed on the current motor drive command in real time to dynamically adjust the control output of the robot actuator.

7. The interactive control method based on multi-source perception as described in claim 1, characterized in that, Based on the constraints and the execution feedback data, security monitoring and response are performed, and effect analysis information is generated based on the execution feedback data to update the optimization model, including: Based on the physical limitations included in the constraints of the working environment, a first-level pressure safety threshold, a second-level pressure safety threshold, and a third-level pressure safety threshold are set for contact pressure, and a first-level trajectory safety threshold, a second-level trajectory safety threshold, and a third-level trajectory safety threshold are set for trajectory deviation. The first-level pressure safety threshold and the first-level trajectory safety threshold correspond to parameter fine-tuning operations, the second-level pressure safety threshold and the second-level trajectory safety threshold correspond to action deceleration operations, and the third-level pressure safety threshold and the third-level trajectory safety threshold correspond to emergency braking operations. The actual contact pressure value and actual pose coordinates in the execution feedback data are acquired in real time. The trajectory deviation between the actual pose coordinates and the motion path in the action strategy is determined. It is determined whether the actual contact pressure value exceeds the first level pressure safety threshold, the second level pressure safety threshold or the third level pressure safety threshold, and whether the trajectory deviation exceeds the first level trajectory safety threshold, the second level trajectory safety threshold or the third level trajectory safety threshold. If the actual contact pressure value or the trajectory deviation exceeds the threshold of the corresponding level, then perform the corresponding parameter fine-tuning operation, action deceleration operation or emergency braking operation according to the threshold level exceeded. After the action strategy is executed, the frequency of pressure exceeding the limit, trajectory smoothness, and internal state recovery time in the execution feedback data are statistically analyzed to generate effect analysis information. The effect analysis information is used as a reward signal for reinforcement learning, and the model parameters of the optimization model are adjusted using the reinforcement learning module.

8. An interactive control device based on multi-source sensing, characterized in that, The multi-source sensing-based interactive control device includes: The multi-source sensing data fusion module is used to collect multi-source heterogeneous sensing data in the working environment and preprocess the multi-source heterogeneous sensing data to generate a standardized sensing dataset. The state and constraint parsing module is used to extract features and semantically associate them with the standardized perception dataset to obtain feature association information, and to parse the current state information of the interactive object and the constraints of the working environment based on the feature association information. The action strategy generation module is used to generate an action strategy containing control parameters and motion path by solving the current state information and the constraints, combined with the task objective, using an optimization model. The execution and feedback acquisition module is used to control the robot execution device to operate according to the action strategy and to collect execution feedback data in real time during the operation. An execution deviation adjustment module is used to determine the execution deviation between the execution feedback data and the control parameters and motion path in the action strategy, and adjust the control output of the robot execution device according to the execution deviation; The safety monitoring and model update module is used to perform safety monitoring and response based on the constraints and the execution feedback data, and to generate effect analysis information based on the execution feedback data to update the optimization model.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and a multi-source perception-based interactive control program stored in the memory and executable on the processor. When executed by the processor, the multi-source perception-based interactive control program implements the steps of the multi-source perception-based interactive control method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores an interactive control program based on multi-source perception, which, when executed by a processor, implements the steps of the interactive control method based on multi-source perception as described in any one of claims 1-7.

Citation Information

Cited By

  • A somatic intelligent multi-modal perception decision method and system

    CN122133090A