Closed-loop control method and device for emotion perception robot
By combining multimodal sensor data processing and deep learning with behavioral decision-making and safety constraints, a closed-loop control method for emotion-aware robots is constructed, which solves the shortcomings of emotion recognition and behavior control, and realizes accurate recognition and reliable control of emotion-aware robots.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN ZHI HUI LIN NETWORK TECH CO LTD
- Filing Date
- 2026-02-13
- Publication Date
- 2026-07-10
AI Technical Summary
Existing methods for controlling emotion-aware robots are inadequate in terms of multimodal perception and feature learning, failing to effectively and accurately identify emotional states. They also lack robust risk assessment mechanisms and execution optimization strategies, resulting in suboptimal control accuracy.
By collecting and preprocessing multimodal sensor data, performing deep feature learning and emotion state vector matching, and combining behavioral decision networks and safety constraint rules to generate control commands, an action planning strategy is constructed. Finally, resource scheduling is performed through an execution optimizer to achieve closed-loop optimization.
It achieves accurate emotion recognition and reliable behavior control, ensuring continuous improvement of control and providing technical support for emotion-aware robots.
Smart Images

Figure CN122363048A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, specifically to a closed-loop control method and device for an emotion-sensing robot. Background Technology
[0002] Existing methods for controlling emotion-aware robots have significant shortcomings. Traditional systems perform poorly in multimodal perception and feature learning, failing to accurately identify emotional states and thus impacting interaction effectiveness.
[0003] Furthermore, existing technologies suffer from bottlenecks in behavioral decision-making and security control. Most systems lack robust risk assessment mechanisms and effective optimization strategies, resulting in suboptimal control precision.
[0004] Existing systems have technical shortcomings in closed-loop optimization. The lack of in-depth analysis of execution states makes it difficult to achieve efficient policy updates through experience-based learning, thus affecting control performance. Solving these problems is crucial for improving the capabilities of emotion-sensing robots. Summary of the Invention
[0005] To address the problems in the existing technology, this application provides a closed-loop control method and device for emotion-sensing robots, which can effectively solve the shortcomings of traditional technologies in emotion recognition, behavior control and strategy optimization, and provide technical support for emotion-sensing robots.
[0006] To solve at least one of the above problems, this application provides the following technical solution: In a first aspect, this application provides a closed-loop control method for an emotion-sensing robot, comprising: The robot's perception system collects visual data streams, voice data streams, and tactile data streams to generate multimodal sensor data. The multimodal sensor data is preprocessed to obtain a standardized feature set. Based on an emotion feature extraction network, the standardized feature set is subjected to deep feature learning to obtain an emotion state vector. The emotion state vector is matched with a preset emotion knowledge base to obtain scene description data. The scene description data is used to perform context association analysis to generate an interaction state vector. The interaction state vector is input into the behavior decision network for strategy planning to obtain behavior control parameters. Based on preset safety constraint rules, the behavior control parameters are risk-assessed to generate safety control indicators. Based on the safety control indicators, an action planning strategy is constructed to obtain an execution sequence graph. The execution sequence graph is time-aligned to generate a control instruction set. The control instruction set is input into the execution optimizer for resource scheduling to obtain a task scheduling matrix. Based on the task scheduling matrix, a robot motion controller is constructed to generate motion control parameters. Based on the motion control parameters, the robot actuator is adjusted in real time to obtain execution status data. The execution status data is input into the memory management unit for experience learning to obtain strategy optimization parameters. Based on the strategy optimization parameters, control update instructions are generated to perform closed-loop optimization of the robot control system.
[0007] Furthermore, it also includes: performing spatial domain filtering on the image data stream acquired by the robot vision sensor to obtain a filtered image set; performing frequency domain transformation on the acoustic signal stream acquired by the speech sensor to obtain an acoustic feature spectrum; sampling and quantizing the pressure signal stream acquired by the tactile sensor to obtain a tactile feature sequence; and using a multimodal synchronization mechanism to timestamp the filtered image set, the acoustic feature spectrum, and the tactile feature sequence to generate multimodal sensor data. The multimodal sensor data is input into a feature preprocessing network for data normalization to obtain a normalized feature set. The normalized feature set is then subjected to noise suppression and dimensionality transformation to obtain a feature dimensionality reduction matrix. Based on a feature selection algorithm, the feature dimensionality reduction matrix is used to filter features and generate a standardized feature set.
[0008] Furthermore, it also includes: inputting a standardized feature set into an emotion feature extraction network to perform hierarchical feature mapping to obtain a hierarchical feature group; performing an attention mechanism to calculate the weight distribution map of the hierarchical feature group; performing weighted fusion of the hierarchical feature group based on the weight distribution map to obtain an emotion state vector; and performing similarity calculation between the emotion state vector and the labeled samples in a preset emotion knowledge base to generate a matching score matrix. A scene semantic parser is constructed based on the matching score matrix to obtain a semantic mapping table. Scene elements are semantically labeled based on the semantic mapping table to generate scene description data. The scene description data is input into a context analysis model to perform temporal association modeling to obtain an interaction state vector.
[0009] Furthermore, it also includes: performing multi-layer progressive analysis on the interaction state vector through a behavior decision network to obtain a decision feature map; prioritizing the decision feature map to generate a task sequence list; performing policy matching on the task sequence list based on a preset behavior template to obtain a candidate policy set; and performing conflict detection on the candidate policy set to generate behavior control parameters. The behavior control parameters are subjected to safety boundary checks to obtain a set of constraint parameters. Based on preset safety constraint rules, the constraint parameter set is subjected to risk measurement to generate a risk score table. The risk score table is then filtered through a safety threshold to obtain safety control indicators.
[0010] Furthermore, it also includes: inputting safety control indicators into the action planner to decompose actions to obtain a primitive action set; combining and reconstructing the primitive action set based on preset action syntax rules to generate an action chain sequence; resolving conflicts in the action chain sequence through a spatiotemporal constraint checker to obtain an execution sequence diagram; and performing time-series marking on the execution sequence diagram to generate a time-series relation table. Based on the timing relationship table, a synchronization controller is constructed to obtain a control timing diagram. The control timing diagram is then decomposed into a task to generate a control instruction set. The control instruction set is then input into the execution optimizer to allocate computational resources and obtain a task scheduling matrix.
[0011] Furthermore, it also includes: inputting the task scheduling matrix into the trajectory planner to generate a path to obtain a motion trajectory map, performing dynamic constraint verification on the motion trajectory map to generate a constraint parameter set, constructing a motion controller based on the constraint parameter set to obtain motion control parameters, and mapping the motion control parameters according to execution priority to generate an execution instruction table; The robot joint actuators are controlled in real time according to the execution instruction table to obtain joint state data. The joint state data is then monitored through a sensor feedback network to generate execution state data. The execution state data is then verified in real time to generate a state feedback signal.
[0012] Furthermore, it also includes: encoding the execution state data using a memory encoder to obtain a state feature group; extracting empirical patterns from the state feature group to generate an empirical sample set; performing reinforcement learning on the empirical sample set based on a preset learning rule to obtain policy optimization parameters; and inputting the policy optimization parameters into an evaluation model to perform performance evaluation and generate an optimization index table. Based on the optimization index table, a feedback compensator is constructed to obtain a set of compensation parameters. The set of compensation parameters is adaptively adjusted to generate control update instructions. The control update instructions are then passed through a closed-loop controller to correct the parameters and obtain a control correction amount. Based on the control correction amount, the robot control system is optimized in a closed loop.
[0013] Secondly, this application provides a closed-loop control device for an emotion-sensing robot, comprising: The interaction analysis module is used to collect visual data streams, voice data streams, and tactile data streams input from the robot's perception system to generate multimodal sensor data. The multimodal sensor data is preprocessed to obtain a standardized feature set. Based on an emotion feature extraction network, deep feature learning is performed on the standardized feature set to obtain an emotion state vector. The emotion state vector is matched with a preset emotion knowledge base to obtain scene description data. Context association analysis is performed based on the scene description data to generate an interaction state vector. The task control module is used to input the interaction state vector into the behavior decision network for strategy planning to obtain behavior control parameters, perform risk assessment on the behavior control parameters based on preset safety constraint rules to generate safety control indicators, construct action planning strategies based on the safety control indicators to obtain an execution sequence graph, perform time sequence alignment processing on the execution sequence graph to generate a control instruction set, and input the control instruction set into the execution optimizer for resource scheduling to obtain a task scheduling matrix. The strategy formulation module is used to construct a robot motion controller based on the task scheduling matrix to generate motion control parameters, to adjust the robot actuator in real time based on the motion control parameters to obtain execution status data, to input the execution status data into the memory management unit for experience learning to obtain strategy optimization parameters, and to generate control update instructions based on the strategy optimization parameters to perform closed-loop optimization of the robot control system.
[0014] Thirdly, this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the closed-loop control method for the emotion-sensing robot.
[0015] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the closed-loop control method for the emotion-sensing robot described above.
[0016] Fifthly, this application provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the closed-loop control method for the emotion-sensing robot.
[0017] As described above, this application provides a closed-loop control method and apparatus for an emotion-perceiving robot. Through multimodal features and deep learning, it achieves accurate emotion recognition. A decision-making mechanism is constructed, combining safety constraints and resource scheduling to establish a reliable execution strategy. Closed-loop optimization is introduced, ensuring continuous improvement of control through experience learning and strategy updates. This method effectively addresses the shortcomings of traditional technologies in emotion recognition, behavior control, and strategy optimization, providing technical support for emotion-perceiving robots. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating the closed-loop control method of the emotion-sensing robot in the embodiments of this application; Figure 2 This is a structural diagram of the closed-loop control device for the emotion-sensing robot in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of the electronic device in the embodiments of this application.
[0020] Figure label: Electronic device 9600, central processing unit 9100, memory 9140, communication module 9110, input unit 9120, audio processor 9130, display 9160, power supply 9170, buffer memory 9141, application / function storage unit 9142, data storage unit 9143, driver storage unit 9144, antenna 9111, speaker 9131, microphone 9132. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0022] The acquisition, storage, use, and processing of data in this application comply with relevant laws and regulations.
[0023] To address the problems existing in current technologies, this application provides a closed-loop control method and apparatus for emotion-aware robots. Through multimodal features and deep learning, it achieves accurate emotion recognition. A decision-making mechanism is constructed, combining safety constraints and resource scheduling to establish a reliable execution strategy. Closed-loop optimization is introduced, ensuring continuous improvement of control through experience learning and policy updates. This method effectively solves the shortcomings of traditional technologies in emotion recognition, behavior control, and policy optimization, providing technical support for emotion-aware robots.
[0024] To effectively address the shortcomings of traditional technologies in emotion recognition, behavior control, and strategy optimization, and to provide technical support for emotion-sensing robots, this application provides an embodiment of a closed-loop control method for emotion-sensing robots. See [link to relevant documentation]. Figure 1 The closed-loop control method for the emotion-sensing robot specifically includes the following: Step S101: Collect visual data stream, voice data stream, and tactile data stream input from the robot's perception system to generate multimodal sensor data. Preprocess the multimodal sensor data to obtain a standardized feature set. Perform deep feature learning on the standardized feature set based on an emotion feature extraction network to obtain an emotion state vector. Perform similarity matching between the emotion state vector and a preset emotion knowledge base to obtain scene description data. Perform context association analysis based on the scene description data to generate an interaction state vector. First, visual, audio, and tactile data streams are input and aligned according to sensor timestamps to generate multimodal sensor data. Using visual frames as reference axes, audio data is aggregated into adjacent visual frames according to a fixed window, and tactile data is backfilled into the same time slice using linear interpolation. Then, cleaning and standardization are performed: spatial filtering and white balance correction are applied to the image data; frequency domain denoising and energy normalization are performed on the audio data; and abnormal impulse removal and amplitude scale alignment are performed on the tactile data. The processed data is encoded into preprocessed blocks with time-slice indices for subsequent model reading.
[0025] Based on the preprocessing block, unified feature engineering is performed on the three types of channels to obtain a standardized feature set. Specifically, visual features are extracted for texture and edge response, speech features for cepstral and fundamental frequency correlation, and tactile features for pressure amplitude and rate of change. Dimensionality reduction mapping is then performed to remove redundant features irrelevant to emotion discrimination, while retaining time slice numbers and channel identifiers. This standardized feature set serves as direct input for emotion modeling and will be continuously used in subsequent matching and context modeling stages.
[0026] The standardized feature set is fed into an emotion feature extraction network (Chinese name: Emotion Mapping Network), where hierarchical convolutions and temporal units are jointly modeled within each channel, and lightweight attention is used to establish connections between channels, outputting an emotion state vector. After reading the channel identifiers, the network automatically reduces the weights of channels with missing or high noise levels to ensure stable output dimensions and positions. The emotion state vector is registered with a time-slice index for subsequent comparison with knowledge base entries.
[0027] Based on the aforementioned emotional state vector, a similarity matching process is initiated to establish a comparison relationship with labeled samples in a pre-defined emotional knowledge base, resulting in a matching score matrix. During matching, multi-scale segments and subject cues are used as retrieval keys. Emotional tags, triggering cues, and occurrence intervals are extracted from high-scoring entries to construct scene description data. The scene description data retains the source time slice and evidence fingerprint, serving as input evidence for downstream contextual analysis.
[0028] The scene description data is read by the interaction context modeler, which models the emotional trajectory and semantic relationships of the same subject in adjacent time slices, and outputs an interaction state vector. The modeler establishes weak links in time, fills in missing segments with placeholders, and adjusts the cross-segment association strength based on cue consistency. The interaction state vector contains subject fingerprints, relationship indicators, and confidence markers, which can be directly consumed by the subsequent decision network.
[0029] In this embodiment, to make the context synthesis process clear and detectable and to facilitate subsequent sorting and risk assessment, a unified one-time weighted expression is used to generate the comprehensive context embedding of the current time slice: Q = a×U + b×V + c×W d×R.
[0030] In the formula, Q is the comprehensive context embedding element of the current time slice, U is the embedding output of the emotion state channel, V is the relationship continuation channel across time slices, W is the scene cue consistency channel, and R is the evidence missing and conflict penalty channel; a, b, c, and d are non-negative weights, which are limited in range during the offline stage and normalized within the sample. After Q is calculated in this segment, it is written into the buffer and read by the output layer of the interaction context modeler in the next segment to form the interaction state vector.
[0031] For example, when visual cues detect a downturned mouth and furrowed brows, accompanied by a rising pitch and increased speech rate, and no significant tactile fluctuations, the U signal primarily originates from the visual and speech sub-channels. The W signal receives a positive gain from the consistency between facial expression and pitch cues. If there is no conflicting evidence, the R signal is low, and ultimately, Q points to a tense emotional region. This result is then written back to the time interval of the scene description data as a weighting basis for weak links across segments.
[0032] The interaction state vector is retained in the output buffer of this step and is read by the behavior decision network as input to the subsequent step S201 to initialize the state of policy planning. Simultaneously, the time interval and cue fields in the scene description data are referenced in step S301 to establish a mapping relationship aligned with downstream features. The standardized feature set and emotion state vector are also allowed to be called by the processing pipeline of step S103 in anomaly backtracking scenarios, forming a closed-loop link from the beginning of this step that allows for evidence location.
[0033] Step S102: Input the interaction state vector into the behavior decision network for strategy planning to obtain behavior control parameters, perform risk assessment on the behavior control parameters based on preset safety constraint rules to generate safety control indicators, construct an action planning strategy based on the safety control indicators to obtain an execution sequence graph, perform time sequence alignment processing on the execution sequence graph to generate a control instruction set, and input the control instruction set into the execution optimizer for resource scheduling to obtain a task scheduling matrix; First, the interaction state vector output in step S101 is used as the initial state input to the behavior decision network (Chinese name: behavior policy generator). The entity fingerprint, relationship indicator, and confidence marker are read sequentially by time slice. After reading, the behavior policy generator filters the set of possible actions based on the current intent context, matches each candidate action with the interaction state vector, and generates a decision feature map. Subsequently, the feature map is sorted and deduplicated according to the task objective and context consistency to form behavior control parameters. These parameters include action type, execution priority, and necessary perception prerequisites, and the time slice index is backfilled for subsequent alignment.
[0034] Based on the behavior control parameters, preset safety constraint rules are loaded to conduct a risk assessment. The assessment process examines the physical boundaries, emotional boundaries, and interaction boundaries item by item, and reads the time interval and cue fields from the scene description data in step S101 to verify whether the current environment meets the conditions for the action to take effect. For items with insufficient evidence or conflicts, a suppression mark is added and an observation and waiting condition is attached. The verified results are summarized into safety control indicators, including risk level, amplitude limit requirements, and stop conditions, and the source mapping key is also retained.
[0035] The safety control indicators are fed into the action planner, which generates an execution sequence diagram in the order of "primitive action - combined action - synchronization condition". The action planner inherits the execution priority from the behavior control parameters, marks the triggering prerequisites and safety limits on each action segment, and inserts intermediate transition nodes for continuous segments that cross positions or objects to avoid direct jumps when evidence is missing. The generated execution sequence diagram uses a main chain to carry the main action and side chains to carry interactive actions, with synchronization points and mutual exclusion relationships registered on the edges.
[0036] Based on the execution sequence diagram, timing alignment is performed. The aligner reads the time slice index and synchronization condition, chains consecutive actions of the same entity into a main chain, and sets mutual exclusion tags for concurrent conflict segments. Missing segments are padded with placeholders, and a reference relationship is established between rollback nodes and the nearest complete checkpoint. After processing, a control instruction set is output. The instruction entries contain action codes, trigger conditions, synchronization points, and safety flags, maintaining a one-to-one mapping with time slices for easy parsing in subsequent scheduling.
[0037] In this embodiment, to facilitate unified consideration in subsequent scheduling after the control instruction set is generated, a one-time weighted expression is used to calculate the scheduling priority of each instruction: T = x×P + y×H + z×K r×M.
[0038] In the formula, T represents the overall priority of the instruction entry, P represents the action utility channel, H represents the context stability channel, K represents the safety importance channel, and M represents the risk penalty channel; x, y, z, and r are non-negative weights, which are limited in range during the offline phase and normalized within the sample. After the calculation of this segment is completed, T is written into the extended field of the control instruction set for direct reading in subsequent resource allocation.
[0039] The control instruction set is input into the execution optimizer, which buckets the instructions according to the size of T and synchronization point constraints, and performs resource scheduling by combining the system-side concurrency limit, backtracking depth limit, and channel characteristics. The execution optimizer generates channel selection and quota allocation for each bucket, and sets delays and retry cycles for instructions with suppression flags. The final output is a task scheduling matrix, where rows correspond to batches, columns correspond to execution channels, and cells record quotas, call order, and threshold adjustment amounts, while retaining evidence mapping keys to support backtracking.
[0040] The task scheduling matrix, as the final product of this step, is read by the trajectory and control construction process in subsequent step S201 to generate motion control parameters. Simultaneously, the synchronization points and safety markers in the control command set are referenced in step S301 to align with newly arrived sensory evidence and trigger backoff when necessary. The above data maintains index consistency with the scene description data in step S101, ensuring that the original time slice can be located for verification in abnormal situations.
[0041] Step S103: Construct a robot motion controller based on the task scheduling matrix to generate motion control parameters, perform real-time adjustment of the robot actuator based on the motion control parameters to obtain execution status data, input the execution status data into the memory management unit for experience learning to obtain strategy optimization parameters, and generate control update instructions based on the strategy optimization parameters to perform closed-loop optimization of the robot control system.
[0042] First, the task scheduling matrix output from step S102 is used as input and parsed into a set of control channels according to two dimensions: batch and channel. The time slice index and evidence mapping key from step S101 are then backfilled. During parsing, the quota, calling order, and threshold adjustment amount are read unit by unit to establish the mapping relationship to three types of control paths: position, speed, and torque, which are used for subsequent controller instantiation.
[0043] Based on the aforementioned control channel set, a trajectory generation and constraint verification chain is constructed. The trajectory generation unit reads the main chain actions from the control command set and splices continuous segments into a path segment sequence according to synchronization points. The dynamic verification unit outputs a set of constraint parameters based on the boundary conditions of joint position, velocity, and acceleration, as well as the amplitude limit and stop requirements of the safety control indicators in step S102. For path segments that do not meet the conditions, segmentation and transition segment insertion are implemented, and the reasons for the transition are recorded for the controller to read.
[0044] The constraint parameter set is read by the motion controller to generate motion control parameters. The controller adopts a hierarchical structure: the upper layer calculates the desired pose and velocity based on the path segment sequence and boundary conditions, and the lower layer distributes three types of sub-control laws—position, velocity, and torque—within each joint servo loop. Each control cycle checks the quota and call order in the task scheduling matrix, postpones over-quota requests to the next cycle, and retains the source batch and time slice index in the instruction header to ensure consistency with the scheduling side.
[0045] Driven by the motion control parameters, the robot's actuators are adjusted in real time, and execution status data is collected. The sensor feedback network transmits joint positions, speeds, drive currents, and external contact signals, while also recording synchronization point attainment and safety flag triggering. Execution status data is written to a status buffer cycle by cycle, bound to path segment numbers and control channel numbers for accurate reference in subsequent learning phases.
[0046] The execution state data is read by the memory management unit, initiating the experience learning process. First, state feature encoding is performed, extracting features such as tracking error, steady-state deviation, and amplitude limiting trigger rate. Then, experience pattern extraction is performed, aggregating similar segments into experience samples and registering their occurrence conditions and action contexts. Based on this, the policy learner is invoked to conduct incremental reinforcement learning, outputting policy optimization parameters, including channel weight adjustments, threshold fine-tuning, and backtracking beat suggestions, along with source batch and time slice indexes for verification.
[0047] Based on the optimized strategy parameters, control update instructions are generated to complete closed-loop optimization. The update process performs fine-grained corrections to the motion controller according to channel type: position path updates path tracking gain and feedforward term, velocity path updates velocity limit and acceleration threshold, and torque path updates torque limit and impedance coefficient. To avoid cross-cycle oscillations, update instructions are versioned and a consistency check is performed through a simulation branch before taking effect; the check record and evidence mapping key are retained together.
[0048] In this embodiment, to uniformly quantify the adoption order and intensity of updates for each channel, a one-time weighted expression is used to calculate the update priority of each control channel: L = p×A + q×B + s×C t×D.
[0049] In the formula, L represents the channel update priority, A represents the error convergence metric channel, B represents the execution stability metric channel, C represents the safety satisfaction metric channel, and D represents the anomaly trigger cost channel; p, q, s, and t are non-negative weights, which are limited in range during the offline phase and normalized within the sample. L is used to arrange the application order of control update instructions; updates above the threshold are executed in the current cycle, and updates below the threshold are merged into the next batch.
[0050] The control update instruction is written back to the robot motion controller at the end of the control cycle to form a new motion control parameter generation strategy. Simultaneously, the version number and change summary are written back to the execution optimizer in step S102 to adjust subsequent quotas and call order. The correspondence between execution status data and strategy optimization parameters is synchronized to the evidence index in step S101, ensuring traceability throughout the entire chain from perception to control to learning, and providing continuous historical constraints for strategy convergence in subsequent cycles.
[0051] As described above, the closed-loop control method for emotion-perceiving robots provided in this application can accurately identify emotions through multimodal features and deep learning. A decision-making mechanism is constructed, combining safety constraints and resource scheduling to establish a reliable execution strategy. Closed-loop optimization is introduced, ensuring continuous improvement of control through experience learning and policy updates. This method effectively addresses the shortcomings of traditional technologies in emotion recognition, behavior control, and policy optimization, providing technical support for emotion-perceiving robots.
[0052] In one embodiment of the closed-loop control method for the emotion-sensing robot of this application, it may further include the following: Step S201: Spatial domain filtering is performed on the image data stream acquired by the robot vision sensor to obtain a filtered image set; frequency domain transformation is performed on the acoustic signal stream acquired by the speech sensor to obtain an acoustic feature spectrum; sampling and quantizing the pressure signal stream acquired by the tactile sensor to obtain a tactile feature sequence; and multimodal sensor data is generated by timestamping the filtered image set, the acoustic feature spectrum, and the tactile feature sequence based on a multimodal synchronization mechanism. Step S202: Input the multimodal sensor data into the feature preprocessing network to normalize the data and obtain a normalized feature group. Perform noise suppression and dimensionality transformation on the normalized feature group to obtain a feature dimensionality reduction matrix. Based on the feature selection algorithm, perform feature filtering on the feature dimensionality reduction matrix to generate a standardized feature set.
[0053] First, the image data stream from the robot's vision sensor is fed into the spatial filtering unit, read frame by frame, and a time-slice index is established. The spatial filtering employs an edge-preserving strategy to suppress high-frequency noise while maintaining the response intensity of corner points and object boundaries, producing a filtered image set. Exposure, white balance, and lens markings are recorded at the frame level for subsequent cross-channel alignment. Simultaneously, the acoustic signal stream from the speech sensor is fed into the frequency domain transformation unit. Frames are segmented and windowed using a window function with fixed length and overlap rate. Amplitude and cepstral correlation are calculated to obtain the acoustic feature spectrum, and frame start and end times are retained. Then, the pressure signal stream from the tactile sensor is sampled and quantized at equal intervals, isolated abnormal pulses are removed, and the rate of change is estimated to obtain the tactile feature sequence.
[0054] Based on the filtered image set, acoustic feature spectrum, and tactile feature sequence, a multimodal synchronization mechanism is invoked for time alignment. This mechanism uses the visual frame as the reference axis, aggregating adjacent acoustic frames to the corresponding visual time slice, and backfilling tactile samples into the same time slice through linear interpolation; gaps across segments are filled with placeholder markers, and alignment quality indicators are recorded. After processing, timestamp markers and channel identifiers are uniformly written to the three types of channels, generating multimodal sensor data, carrying the source device and frame index, for downstream networks to read and trace back.
[0055] The multimodal sensor data is input into a feature preprocessing network for normalization. Normalization is performed within each channel: on the visual side, range remapping and white balance constraints are applied to brightness and color distribution; on the speech side, range alignment is applied to energy and cepstral coefficients; and on the tactile side, scale alignment is applied to pressure amplitude and rate of change. The normalization results are organized into normalized feature groups by time slice, and alignment quality indicators and channel identifiers are retained at the entry level to ensure that subsequent suppression strategies are effective slice-by-slice.
[0056] Based on the normalized feature set, noise suppression and dimensionality transformation are performed. In the noise suppression stage, alignment quality indicators are read, low-quality segments have their channel weights reduced and smoothed to avoid single-channel anomalies interfering with the overall representation. Dimensionality transformation employs a two-stage process of intra-channel mapping and cross-channel aggregation, compressing redundant correlations while maintaining a one-to-one correspondence between time slices, and outputting a feature dimensionality reduction matrix. This matrix is organized with time slices as rows and channel aggregation axes as columns, retaining reverse pointers mapped to the original timestamps.
[0057] The reduced feature matrix is fed into a feature selection algorithm for filtering, generating a standardized feature set. During the filtering phase, based on a pre-defined feature mask related to the task objective and emotion, and combined with a channel stability index, items with low contribution to emotion estimation and instability are removed, while key components sensitive to facial expression changes, voice quality, and tactile changes are retained. The standardized feature set carries a time-slice index, channel identifier, and filtering mask number at the sample level to ensure repeatable extraction and auditability.
[0058] To facilitate a unified consideration in subsequent behavioral decision-making stages, this embodiment uses a one-time weighted representation for multi-channel features within the same time slice to form a comprehensive reference value before screening and control the intensity of the penalty: N = g×S + h×E + j×C k×Z.
[0059] In the formula, N is the comprehensive reference element, S is the visual feature channel, E is the speech feature channel, C is the tactile feature channel, and Z is the joint penalty channel for alignment quality and noise; g, h, j, and k are non-negative weights, which are limited in range during the offline stage and normalized within the samples. The N does not directly replace the screening result, but is only used to adjust the candidate threshold during the screening process. The final output is still based on the standardized feature set obtained by screening the feature dimensionality reduction matrix.
[0060] The standardized feature set is directly read as input to the emotion feature extraction network in step S101, and used for hierarchical mapping and attention weighting to generate an emotion state vector. Simultaneously, the timestamps and channel identifiers in the multimodal sensor data are referenced in step S102 to align and backtrack the control command set with perceptual evidence. The alignment quality indicator and reverse pointer previously registered in the synchronization mechanism are invoked by the control side in step S103 to locate the corresponding time slice and trigger a rollback in case of execution anomalies. Through this connection, steps S201 and S202 complete the processing link from raw acquisition to learnable representation, and provide traceable time and channel indices for subsequent stages.
[0061] In one embodiment of the closed-loop control method for the emotion-sensing robot of this application, it may further include the following: Step S301: Input the standardized feature set into the emotion feature extraction network to perform hierarchical feature mapping to obtain hierarchical feature groups, perform attention mechanism calculation on the hierarchical feature groups to generate a weight distribution map, perform weighted fusion on the hierarchical feature groups based on the weight distribution map to obtain an emotion state vector, and perform similarity calculation between the emotion state vector and the labeled samples in the preset emotion knowledge base to generate a matching score matrix. Step S302: Construct a scene semantic parser based on the matching score matrix to obtain a semantic mapping table, perform semantic annotation on scene elements based on the semantic mapping table to generate scene description data, and input the scene description data into the context analysis model to perform temporal association modeling to obtain an interaction state vector.
[0062] First, the standardized feature set output from step S202 is read into the emotion feature extraction network (Chinese name: Emotion Mapping Network) by time slice, and hierarchical feature mapping is completed within each channel. The mapping process extracts local patterns independently from three branches: visual, speech, and tactile, generating edge texture units, acoustic cepstral units, and pressure change units respectively, and then merging them into hierarchical feature groups at each time slice. To ensure consistency with the previous synchronization relationship, the time slice index and channel identifier are retained within the hierarchical feature group, which can be used for subsequent slice-by-slice weighting.
[0063] Based on the hierarchical feature groups, an attention mechanism is initiated for computation. The attention unit first generates saliency suggestions based on channel stability and cue consistency, and then generates a weight distribution map by combining missing markers within the time slice. This weight distribution map simultaneously covers both the channel dimension and the sub-feature dimension, explicitly downsizing components with lower alignment quality to avoid abnormal amplification of single channels. Subsequently, the hierarchical feature groups are weighted and fused according to the weight distribution map, outputting an emotion state vector, and the time slice index is written back to the vector entries to ensure a one-to-one correspondence with subsequent knowledge base comparisons.
[0064] The emotional state vector is fed into a similarity calculation unit and compared one by one with labeled samples in a pre-defined emotional knowledge base to obtain a matching score matrix. During the comparison process, only the index consistent with the time slice is used for retrieval to prevent cross-segment confusion; simultaneously, triggering clues and subject identifiers are recorded to facilitate subsequent semantic analysis and evidence retrieval. The matching score matrix retains candidate sample numbers and similarity scores, and sets candidate labels for low-confidence samples, which are not yet involved in the main path determination.
[0065] Based on the matching score matrix, a scene semantic parser is constructed to obtain a semantic mapping table. The semantic parser reads the candidate sample number and triggering clues, mapping general emotion tags to a set of scene elements, including three categories of items: subject, object, and environmental factors. For multiple candidates with similar confidence levels, hierarchical labeling is given based on clue consistency and contextual priors, and the second highest-scoring item is included in the candidate layer. The semantic mapping table is organized using time slices as keys, preserving the index correspondence with the emotion state vector.
[0066] The semantic mapping table is used to semantically annotate scene elements and output scene description data. The annotation process writes the subject fingerprint, emotion tag, and triggering clues into the same entry, forming explicit relationship pairs for cross-subject interactions; gaps in time slices are filled with placeholder markers, and the reason for the gap and the reference segment are recorded. The scene description data includes the source time slice and evidence fingerprint at the entry level, which can be used for consistency verification in subsequent steps.
[0067] Based on the scenario description data, it is input into a context analysis model (Chinese name: Interaction Context Modeler) for temporal association modeling. This modeler reads the subject fingerprints and relationship pairs, constructs a continuous chain along the timeline, and sets weak associations for short-term inconsistent entries to prevent premature chain breakage; it also establishes switching nodes at time points where environmental factors change abruptly to ensure clear boundaries between preceding and following contexts. After processing, it outputs an interaction state vector containing three types of fields: subject chain identifier, relationship indicator, and confidence marker.
[0068] Based on the aforementioned interaction state vector, a closed-loop reference to the upstream product is further established on the model output side. Specifically, the interaction state vector retains a reverse pointer to the scene description data, facilitating direct indexing of emotion tags and trigger cues during strategy planning in subsequent step S102; simultaneously, it retains the mapping between the emotion state vector and the matching score matrix, allowing step S103 to backtrack to the corresponding time slice and verify the channel weights and candidate status at that time in the event of an execution anomaly. The aforementioned reference relationships are registered together when written to the output buffer to avoid information loss across stages.
[0069] For example, when the visual branch continuously identifies a furrowed brow in adjacent time slices, the speech branch shows a short-term increase in intensity while the fundamental frequency remains stable, and the tactile branch shows no significant fluctuations, the weight distribution map will increase the weights of visual and some speech sub-features. The weighted fusion of these emotional state vectors will more closely resemble the labeled sample of "nervous but cooperative." Based on this, the semantic mapping table generates a set of elements with the user as the subject, the robot as the object, and the indoor environment as the setting. In the interaction context modeling, the relationship of "needing reassuring feedback" is attached to the subject chain, forming an interaction state vector that can be directly used for policy planning.
[0070] Finally, the interaction state vector is read by the behavior decision network as the input item in step S102 to generate behavior control parameters; the scene description data and the matching score matrix are referenced in the risk assessment stage of step S102 to verify the action premise and safety boundary; the emotion state vector is used in the anomaly backtracking stage of step S103 to verify the feature fusion and weighting basis at that time, forming a traceable chain from feature learning, knowledge matching to context modeling.
[0071] In one embodiment of the closed-loop control method for the emotion-sensing robot of this application, it may further include the following: Step S401: The interaction state vector is analyzed in multiple layers through a behavior decision network to obtain a decision feature map. The decision feature map is prioritized to generate a task sequence list. Based on a preset behavior template, the task sequence list is matched with a strategy to obtain a candidate strategy set. The candidate strategy set is subjected to conflict detection to generate behavior control parameters. Step S402: Perform a safety boundary check on the behavior control parameters to obtain a set of constraint parameters, perform risk measurement on the set of constraint parameters based on preset safety constraint rules to generate a risk score table, and filter the risk score table through a safety threshold to obtain a safety control index.
[0072] First, the interaction state vector output in step S301 is input into the behavior decision network (Chinese name: behavior strategy generator) in time-slice order. The main chain identifier, relationship indicator, and confidence marker are read and associated with the time interval and triggering clues in the scene description data from step S301. Within each time slice, the behavior strategy generator filters out a set of actionable actions based on the intent context. Each action is compared with the interaction state vector, calculating the consistency between the action and emotion / context, and aggregating them into a decision feature map. Nodes in this feature map represent candidate actions, edges represent sequential relationships and synchronization dependencies, and the source time slice and evidence fingerprint are retained on the nodes for subsequent backtracking.
[0073] Based on the decision feature map, priority sorting is performed to generate a task sequence list. The sorter reads the consistency score, context continuity, and confidence flag of the interaction state vector on the nodes, reduces the sorting weight of time slices with gaps, and deduplicates and merges similar actions to form a relatively concise sequence. The sorting result provides one or a few candidate actions for each time slice, arranged in descending order of priority; at the same time, reference relationships are preserved for entries that have dependencies across time slices to prevent chain breaks in subsequent matching stages.
[0074] The task sequence list is sent to the strategy matching unit, which establishes a mapping relationship with the preset behavior templates to obtain a candidate strategy set. The behavior template is expressed in the structure of "trigger premise - action combination - synchronization point - fallback condition". During matching, the emotion tag and trigger clues from step S301 are read to expand a single action into an action chain that satisfies the current context. For multiple templates that are matched at the same time, the matching unit assigns different matching levels according to the continuity and relationship indication of the main chain, registers the suboptimal template as a candidate item, and carries the required perceptual premise and execution constraint on the item.
[0075] Based on the candidate strategy set, conflict detection is performed, and behavioral control parameters are output. The detector identifies three types of problems: resource mutual exclusion, synchronization conflict, and semantic contradiction. Resource mutual exclusion is identified by comparing the occupancy of the same control channel within the same time slice; synchronization conflict is identified by checking whether the synchronization points of template edges overlap and have no common triggering premise; semantic contradiction is identified by comparing the compatibility between the relationship indication and the action target. Conflict entries are pruned, transitional actions are inserted, or mutual exclusion relationships are set. The generated behavioral control parameters include action type, execution priority, triggering premise, and synchronization point description, and retain the source index for backtracking to the task sequence table and the decision feature map in case of anomalies.
[0076] Based on the behavior control parameters, the system enters the safety boundary verification stage, obtaining a set of constraint parameters. The safety verifier reads three types of rules item by item: physical boundaries, emotional boundaries, and interaction boundaries. Physical boundaries are derived from the motion and mechanical limitations of the actuator; emotional boundaries are derived from the emotion labels and intensity in the scene description data; and interaction boundaries are derived from constraints such as dialogue turns and human-computer distance. On the one hand, the verifier imposes limits and distance buffers on the amplitude, speed, and action distance of actions; on the other hand, it applies delays or observation conditions to items involving sensitive words or actions. The output set of constraint parameters is bound to the original parameters item by item.
[0077] The constraint parameter set is input into the risk measurement unit to generate a risk scoring table. The measurement unit aggregates three types of indicators: first, the cost of occurrence, referring to the probability of potential collisions, boundary crossings, and accidental injury; second, uncertainty, referring to the confidence markers and time slice gaps in the interaction state vector; and third, situational sensitivity, referring to the strength of emotion tags and triggering cues. Each item provides a risk level and suggested action boundaries, and records the rule number and evidence source used to ensure traceability in subsequent audits.
[0078] Based on the risk scoring table, safety threshold filtering is performed to obtain safety control indicators. The filter suppresses high-risk items according to system preset threshold conditions, adds limiting, speed reduction, and stopping conditions to medium-risk items, and retains the original priority for low-risk items. To ensure continuity with subsequent step S102, each item in the safety control indicators retains the execution priority, synchronization point, and trigger prerequisite fields, and the filtering results are written back to the behavior control parameters in a tagged form to ensure that no recalculation is needed when entering action planning and resource scheduling.
[0079] Based on the aforementioned safety control indicators, step S102 can directly read the action boundaries and risk levels, perform action decomposition and timing alignment, and ultimately generate a control instruction set. Simultaneously, the risk scoring table serves as a reference in the resource scheduling of step S102, determining whether instructions with suppression markers should be delayed or retried. If an abnormal event is detected on the execution side in subsequent step S103, the source index of the behavior control parameters can be used to trace back to the task sequence table and the decision feature map. The judgment criteria at the time can be verified by combining the scene description data and the interaction state vector, forming a closed-loop chain of evidence.
[0080] In one embodiment of the closed-loop control method for the emotion-sensing robot of this application, it may further include the following: Step S501: Input safety control indicators into the action planner to decompose actions and obtain primitive action sets. Based on preset action syntax rules, combine and reconstruct the primitive action sets to generate action chain sequences. Use the spatiotemporal constraint checker to resolve conflicts in the action chain sequences to obtain an execution sequence diagram. Perform time-series marking on the execution sequence diagram to generate a time-series relation table. Step S502: Construct a synchronization controller based on the timing relationship table to obtain a control timing diagram, decompose the control timing diagram into a task to generate a control instruction set, and input the control instruction set into the execution optimizer to allocate computational resources to obtain a task scheduling matrix.
[0081] First, the safety control indicators output in step S402 are input into the action planner. The execution priority, synchronization point, and triggering prerequisites for each entry are read, and associated with the behavior control parameter source index from step S401. Based on this, the action planner performs action decomposition, breaking down complex actions into executable primitive action sets, and registering the required perception prerequisites, spatial pose, and duration range for each action. During the decomposition process, entries with conditions such as amplitude limiting, deceleration, or stopping are expanded with independent constraint fields at the primitive level to prevent constraint loss during subsequent combination.
[0082] Based on the set of primitive actions, actions are combined and reconstructed according to preset action syntax rules to generate an action chain sequence. The action syntax rules are organized in a "start-transition-terminate" structure, requiring adjacent primitives to be connectable in terms of pose, channel, and perception premises. For cases involving alternating subjects or cross-object operations, bridging primitives are introduced as transitions to ensure continuity of constraints within the chain. The action chain sequence retains references to safety control indicators at the node level and maps the triggering premises in step S401 to start and stop conditions on the chain.
[0083] The action chain sequence is fed into a spatiotemporal constraint checker for conflict resolution, outputting an execution sequence graph. The checker performs mutual exclusion checks on resource usage within the same time slice, consistency checks on synchronization points across time slices, and adds buffers to spatial segments that may cause collisions or out-of-bounds errors. If concurrent conflicts are detected, waiting nodes are inserted into lower-priority branches; if spatial conflicts are detected, the sequence is split into sequentially executed sub-segments and marked with transitions. The processed execution sequence graph uses a main chain to carry continuous operations, side chains to record interactive or auxiliary operations, and retains triggering prerequisites and constraints on the edges.
[0084] Based on the execution sequence diagram, timing markers are applied to generate a timing relationship table. The marker reads the start / stop conditions and synchronization points within the chain, registering parallelizable segments, serializable segments, and rollback checkpoints respectively; for entries involving observation and waiting, the minimum waiting interval and maximum delay threshold are recorded. The timing relationship table is organized by time slice index, with reference keys for the source action chain number and safety control indicators, facilitating direct parsing by the subsequent synchronization controller.
[0085] The timing relationship table is used to construct the synchronization controller, resulting in a control timing diagram. The synchronization controller is structured in three layers: "time axis - main chain - relation side chain". Resource quota boundaries are applied to parallel segments within the same main chain, and trigger dependencies are established for collaborative segments across main entities. The control timing diagram labels each node with its expected arrival time, latest execution time, and synchronization condition satisfaction flag, and establishes bidirectional references between rollback nodes and the nearest checkpoint to support rapid recovery from the running state.
[0086] Based on the control timing diagram, task decomposition is performed to generate a control instruction set. The task decomposer maps high-level action nodes to instruction entries for low-level control channels, including three types of parameter combinations: position target, speed limit, and torque boundary, and inherits the start / stop conditions and delay thresholds from the timing diagram. For entries with observation wait, a perception readback condition is added; for entries with safety limits, the amplitude and rate limits are explicitly applied to the instruction parameters. The generated control instruction set maintains a one-to-one mapping with time slices and retains reverse pointers referencing the execution sequence diagram and timing relationship table.
[0087] The control instruction set is input into the execution optimizer to allocate computational resources, resulting in a task scheduling matrix. The execution optimizer first buckets the instructions based on the characteristics of the control channels and the system's concurrency limit, prioritizing main chain instructions requiring low latency response to fast channels, and allocating side chain and wait / observe instructions to regular channels. Then, based on the latency threshold and the latest execution time, it assigns quotas and call order to each bucket, and records the retry cycle time and backtracking depth. The final output task scheduling matrix is arranged with batches as rows and channels as columns, with each cell recording quota, order, and threshold adjustment amounts, and carrying an index key referencing the control instruction set.
[0088] Simultaneously with writing the output to the buffer, the version identifiers of the execution sequence diagram and timing relationship table are written back to the security control indicator in step S402 for subsequent tracking of scheduling changes under the same security assumption. The start / stop conditions and observation / wait conditions in the control instruction set are also exposed to the resource scheduling loop in step S102, allowing for adjustments to priority and concurrency strategies based on new evidence during operation. Through this connection, steps S501 and S502 consistently apply security-side constraints and timing-side organization to the schedulable instruction layer, forming a direct input to the subsequent step S103.
[0089] In one embodiment of the closed-loop control method for the emotion-sensing robot of this application, it may further include the following: Step S601: Input the task scheduling matrix into the trajectory planner to generate a path and obtain a motion trajectory map. Perform dynamic constraint verification on the motion trajectory map to generate a constraint parameter set. Construct a motion controller based on the constraint parameter set to obtain motion control parameters. Map the motion control parameters according to execution priority to generate an execution instruction table. Step S602: Perform real-time control of the robot joint actuator according to the execution instruction table to obtain joint state data, use the joint state data to perform state monitoring through the sensor feedback network to generate execution state data, and perform real-time verification of the execution state data to generate state feedback signals.
[0090] First, the task scheduling matrix output from step S502 is input into the trajectory planner by batch and channel dimensions. The quota, calling order, delay threshold, and backtracking depth of each unit are read, and associated with the control instruction set index key from step S501. Based on three types of parameters—position target, speed limit, and torque boundary—the trajectory planner generates continuous path segments on the main chain and asynchronous branches on the side chains. For entries with observation waiting conditions, placeholder segments are first generated, and then expanded into executable segments after the perception backtracking is satisfied. After processing, the results are synthesized into a motion trajectory diagram, and start / stop conditions and synchronization points are registered on the edges.
[0091] Based on the motion trajectory diagram, a dynamic constraint check is performed, and a constraint parameter set is output. The check unit performs boundary checks on joint positions, velocities, accelerations, and torques for each path segment, and supplements safety margins by referring to the amplitude limit and stop conditions in the safety control indicators of step S402. For path segments with potential collision or boundary crossing risks, buffer zones and maximum transition times are marked; for path segments that do not meet the constraints, they are split into sub-segments and smooth transition segments are inserted, while recording the reasons for the splitting and the corresponding scheduling matrix unit positions to ensure that subsequent controllers can locate the source.
[0092] The constraint parameter set is used to construct the motion controller and generate motion control parameters. The controller adopts a hierarchical structure: the upper layer is a path tracking loop that receives path segments and start / stop conditions, and calculates the desired pose and velocity; the lower layer is a joint servo loop that generates underlying control quantities for the position, velocity, and torque control channels, respectively. To maintain consistency with the scheduling side, the controller checks the quota and call order of its batch in each control cycle, delays requests exceeding the quota, and records the reason for the delay to avoid jitter caused by resource preemption. The generated motion control parameters carry time slice indexes and path segment numbers for consistent referencing during the execution and feedback phases.
[0093] Based on the motion control parameters, a hierarchical mapping is performed to form an execution instruction table. The hierarchical mapping is based on the priority from the task scheduling matrix and the synchronization conditions of step S501. High-priority main chain entries are assigned compact timing and low-latency targets, while side chain and observation / waiting entries are set with relaxed timing and readback thresholds. Each instruction includes the target pose, upper velocity limit, and torque boundary, and retains start / stop conditions, the latest execution time, and a rollback checkpoint index for rapid judgment and rollback during the execution phase.
[0094] The execution instruction table is used to control the robot's joint actuators in real time, obtaining joint state data. The controller issues position, velocity, and torque commands according to the instruction cycle. The actuators return raw quantities such as joint angle, joint velocity, drive current, and temperature at a uniform sampling period, and immediately generate event flags when a synchronization point is reached or a safety flag is triggered. The joint state data is written to the running buffer cycle by cycle, maintaining a one-to-one correspondence with the entry level of the instruction table, ensuring that subsequent monitoring and verification can locate specific instructions.
[0095] Based on the joint status data, a sensor feedback network is used for status monitoring to generate execution status data. The monitoring process de-jitters key quantities and removes outliers, and determines whether timeouts or non-compliance have occurred based on start / stop conditions and the latest execution time. For path segments with buffers, proximity and contact trends are additionally calculated for early warning. The execution status data includes achievement markers, error statistics, and event sequences, organized by time slice and path segment number index for easy backtracking and learning.
[0096] The execution status data is read by the real-time verification unit, generating a status feedback signal. Verification is based on three types of rules: first, a trajectory tracking rule, comparing the deviation between the expected pose and the measured pose to see if it is within the allowable range; second, a timing consistency rule, checking if the synchronization point is reached before the latest execution time; and third, a safety boundary rule, verifying whether the torque and speed exceed the limits. If any rule is not met, a degradation suggestion and a rollback checkpoint index are included in the feedback signal, indicating the need to trigger observation waiting or stop waiting; if all rules are met, it is marked as a state where the cycle time can be relaxed further.
[0097] Based on the aforementioned status feedback signals, step S103 can directly read the data for experience learning and strategy updates; simultaneously, the correspondence between execution status data and joint status data is written back to the task scheduling matrix index in step S502 for fine-tuning resource allocation in the next cycle. To maintain consistency, the version identifiers of the motion trajectory diagram and constraint parameter set are archived along with the execution instruction table, facilitating the reconstruction of the control context during anomaly replay. Through the above connection, steps S601 and S602 integrate scheduling, path, control, and feedback into an executable closed loop, providing a stable entry point for subsequent learning and parameter injection.
[0098] In one embodiment of the closed-loop control method for the emotion-sensing robot of this application, it may further include the following: Step S701: Encode the execution state data using a memory encoder to obtain a state feature group, extract the state feature group using an experience pattern to generate an experience sample set, perform reinforcement learning on the experience sample set based on a preset learning rule to obtain policy optimization parameters, and input the policy optimization parameters into an evaluation model to perform performance evaluation and generate an optimization index table. Step S702: Construct a feedback compensator based on the optimization index table to obtain a compensation parameter set, perform adaptive adjustment on the compensation parameter set to generate a control update command, perform parameter correction on the control update command through a closed-loop controller to obtain a control correction amount, and perform closed-loop optimization on the robot control system based on the control correction amount.
[0099] First, the execution status data output in step S602 is input into the memory encoder according to the time slice and path segment number, and three types of information are read: achievement flag, error statistics, and event sequence. The memory encoder completes feature encoding at the item level, mapping tracking deviation, steady-state deviation, synchronization point advance or lag, amplitude limiting trigger rate, and backoff frequency to state feature groups; at the same time, it retains the reference keys of source batch, control channel, and start / stop conditions to ensure that subsequent learning stages can distinguish and process according to scenario and task type.
[0100] Based on the state feature set, empirical patterns are extracted to form an empirical sample set. The extraction process uses time windows as the granularity, aggregating similar trajectories within adjacent windows to output several empirical segments; segments containing anomalous events are separated into separate volumes to avoid dilution by normal samples. Each empirical sample carries a triggering context, a control instruction segment, and a corresponding execution outcome label, and registers the scheduling conditions and safety limit settings that generated the sample, facilitating the establishment of causal relationships during the learning phase.
[0101] The experience sample set is fed into a policy learner, where reinforcement learning is performed according to preset learning rules, outputting policy optimization parameters. The learner is organized in a three-element structure of "state-action-result". The state is obtained by concatenating state feature groups with context constraints, the action corresponds to the fine-tuning of channel weights and thresholds within the controller, and the result comes from the joint signal of execution achievement and safety event. To avoid overfitting to a single scenario, the learner performs sampling equalization on samples from the same source and with similar results, and sets a decay coefficient for samples with backsliding to ensure stable update direction.
[0102] Based on the optimized strategy parameters, they are input into the evaluation model for offline performance evaluation, generating an optimization index table. The evaluation model calculates three types of indicators: response timing consistency, trajectory error convergence, and safety boundary satisfaction rate. Samples containing abnormal events are scored separately to avoid the average value masking risk points. The optimization index table provides suggested adjustment levels for each control channel, while retaining evidence citation keys and time slice ranges for easy comparison and verification with the original execution status data.
[0103] The optimized index table was used to construct the feedback compensator, resulting in a set of compensation parameters. The compensator generates three types of compensation quantities on a channel-by-channel basis: position feedforward, velocity limiting correction, and torque impedance adjustment. For scenarios with timing lag, a small-amplitude phase lead is introduced at the beat layer. For scenarios that frequently trigger limiting, the target boundary is appropriately tightened and the sensitivity of backoff triggering is increased. Each compensation parameter is bound to its corresponding control channel and task type to ensure a clear application scope.
[0104] Based on the compensation parameter set, adaptive adjustment is performed, and control update instructions are output. The adaptive module reads the sliding statistics of recent execution status data, sets upper limits and change rate limits on the compensation amount to prevent cross-cycle oscillations; for channels with large fluctuations, the application is delayed and additional observation windows are required to reduce the risk of misjudgment. The generated control update instructions include the update target, amplitude level, and effective conditions, and retain the reverse pointer to the optimization index table.
[0105] The control update command is calibrated using a closed-loop controller to obtain the control correction value. The closed-loop controller first applies the command in the simulation branch, verifying whether the evidence reference key matches the current scheduling conditions. If they do not match, the system degrades to observation mode. Once a match is found, the command is written to the shadow configuration of the online branch and takes effect sequentially in the next control cycle according to priority. The version number and scope of impact are recorded during application of the control correction value for easy backtracking later.
[0106] Based on the aforementioned control corrections, closed-loop optimization is performed on the robot control system. The optimization action is triggered at the control cycle boundary, with the position, velocity, and torque channels updated sequentially. At the start of the next sampling cycle, the actuator side transmits the first batch of joint state data. If the state feedback indicates that the error and risk indicators are changing in the target direction, the update cycle is maintained; if it deviates, a rollback to the previous version is immediately triggered, and new experience samples are added to the memory encoder side for use in the next learning round. Through this round-trip link, steps S701 and S702 precipitate the execution data as strategy revisions, which are then injected back into the controller in a controlled manner, forming a traceable and rollback-capable closed-loop update path.
[0107] To effectively address the shortcomings of traditional technologies in emotion recognition, behavior control, and strategy optimization, and to provide technical support for emotion-sensing robots, this application provides an embodiment of a closed-loop control device for an emotion-sensing robot, which implements all or part of the closed-loop control method of the aforementioned emotion-sensing robot. See [link to embodiment]. Figure 2 The closed-loop control device for the emotion-sensing robot specifically includes the following components: The interaction analysis module 10 is used to collect visual data streams, voice data streams, and tactile data streams input from the robot's perception system to generate multimodal sensor data, preprocess the multimodal sensor data to obtain a standardized feature set, perform deep feature learning on the standardized feature set based on an emotion feature extraction network to obtain an emotion state vector, perform similarity matching between the emotion state vector and a preset emotion knowledge base to obtain scene description data, and perform context association analysis based on the scene description data to generate an interaction state vector. Task control module 20 is used to input the interaction state vector into the behavior decision network for strategy planning to obtain behavior control parameters, perform risk assessment on the behavior control parameters based on preset safety constraint rules to generate safety control indicators, construct action planning strategies based on the safety control indicators to obtain an execution sequence diagram, perform time sequence alignment processing on the execution sequence diagram to generate a control instruction set, and input the control instruction set into the execution optimizer for resource scheduling to obtain a task scheduling matrix. The strategy formulation module 30 is used to construct a robot motion controller based on the task scheduling matrix to generate motion control parameters, to perform real-time adjustment of the robot actuator based on the motion control parameters to obtain execution status data, to input the execution status data into the memory management unit for experience learning to obtain strategy optimization parameters, and to generate control update instructions based on the strategy optimization parameters to perform closed-loop optimization of the robot control system.
[0108] As described above, the closed-loop control device for the emotion-perceiving robot provided in this application embodiment can accurately identify emotions through multimodal features and deep learning. A decision-making mechanism is constructed, combining safety constraints and resource scheduling to establish a reliable execution strategy. Closed-loop optimization is introduced, ensuring continuous improvement of control through experience learning and strategy updates. This method effectively addresses the shortcomings of traditional technologies in emotion recognition, behavior control, and strategy optimization, providing technical support for emotion-perceiving robots.
[0109] From a hardware perspective, in order to effectively address the shortcomings of traditional technologies in emotion recognition, behavior control, and strategy optimization, and to provide technical support for emotion-sensing robots, this application provides an embodiment of an electronic device for implementing all or part of the closed-loop control method of the emotion-sensing robot. The electronic device specifically includes the following components: The system comprises a processor, memory, a communications interface, and a bus; wherein the processor, memory, and communications interface communicate with each other via the bus; the communications interface is used to realize information transmission between the closed-loop control device of the emotion-sensing robot and core business systems, user terminals, and related databases and other related devices; the logic controller can be a desktop computer, tablet computer, or mobile terminal, etc., and this embodiment is not limited to these. In this embodiment, the logic controller can be implemented with reference to the embodiments of the closed-loop control method and the closed-loop control device of the emotion-sensing robot in the embodiments, the contents of which are incorporated herein, and repeated details will not be described again.
[0110] It is understood that the user terminal may include smartphones, tablet computers, network set-top boxes, portable computers, desktop computers, personal digital assistants (PDAs), in-vehicle devices, smart wearable devices, etc. Among these, the smart wearable devices may include smart glasses, smartwatches, smart bracelets, etc.
[0111] In practical applications, the closed-loop control method of the emotion-sensing robot can be partially executed on the electronic device side as described above, or all operations can be completed in the client device. The choice can be made based on the processing power of the client device and the limitations of the user's usage scenario. This application does not impose any limitations on this. If all operations are completed in the client device, the client device may further include a processor.
[0112] The aforementioned client device may have a communication module (i.e., a communication unit) that can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side; in other implementation scenarios, it may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, a server cluster consisting of multiple servers, or a distributed server structure.
[0113] Figure 3 This is a schematic block diagram illustrating the system configuration of the electronic device 9600 according to an embodiment of this application. Figure 3 As shown, the electronic device 9600 may include a central processing unit 9100 and a memory 9140; the memory 9140 is coupled to the central processing unit 9100. It is worth noting that... Figure 3 This is an example; other types of structures can also be used to supplement or replace this structure to achieve telecommunications functions or other functions.
[0114] In one embodiment, the closed-loop control method functionality of the emotion-aware robot can be integrated into the central processing unit 9100. The central processing unit 9100 can be configured to perform the following control: Step S101: Collect visual data stream, voice data stream, and tactile data stream input from the robot's perception system to generate multimodal sensor data. Preprocess the multimodal sensor data to obtain a standardized feature set. Perform deep feature learning on the standardized feature set based on an emotion feature extraction network to obtain an emotion state vector. Perform similarity matching between the emotion state vector and a preset emotion knowledge base to obtain scene description data. Perform context association analysis based on the scene description data to generate an interaction state vector. Step S102: Input the interaction state vector into the behavior decision network for strategy planning to obtain behavior control parameters, perform risk assessment on the behavior control parameters based on preset safety constraint rules to generate safety control indicators, construct an action planning strategy based on the safety control indicators to obtain an execution sequence graph, perform time sequence alignment processing on the execution sequence graph to generate a control instruction set, and input the control instruction set into the execution optimizer for resource scheduling to obtain a task scheduling matrix; Step S103: Construct a robot motion controller based on the task scheduling matrix to generate motion control parameters, perform real-time adjustment of the robot actuator based on the motion control parameters to obtain execution status data, input the execution status data into the memory management unit for experience learning to obtain strategy optimization parameters, and generate control update instructions based on the strategy optimization parameters to perform closed-loop optimization of the robot control system.
[0115] As described above, the electronic device provided in this application embodiment achieves accurate emotion recognition through multimodal features and deep learning. A decision-making mechanism is constructed, combining safety constraints and resource scheduling to establish a reliable execution strategy. Closed-loop optimization is introduced, ensuring continuous improvement of control through experience learning and policy updates. This method effectively addresses the shortcomings of traditional technologies in emotion recognition, behavior control, and policy optimization, providing technical support for emotion-sensing robots.
[0116] In another embodiment, the closed-loop control device of the emotion-sensing robot can be configured separately from the central processing unit 9100. For example, the closed-loop control device of the emotion-sensing robot can be configured as a chip connected to the central processing unit 9100, and the closed-loop control method function of the emotion-sensing robot can be realized through the control of the central processing unit.
[0117] like Figure 3 As shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It is worth noting that the electronic device 9600 does not necessarily need to include these components. Figure 3 All components shown; in addition, the electronic device 9600 may also include Figure 3 For components not shown, please refer to existing technologies.
[0118] like Figure 3 As shown, the central processing unit 9100, sometimes also referred to as a controller or operating control, may include a microprocessor or other processor device and / or logic device, which receives inputs and controls the operation of various components of the electronic device 9600.
[0119] The memory 9140 may be, for example, one or more of a cache, flash memory, hard drive, removable media, volatile memory, non-volatile memory, or other suitable devices. It may store the aforementioned failure-related information, and also store a program for executing that information. The central processing unit 9100 may execute the program stored in the memory 9140 to perform information storage or processing, etc.
[0120] Input unit 9120 provides input to central processing unit 9100. Input unit 9120 may be, for example, a keypad or touch input device. Power supply 9170 provides power to electronic device 9600. Display 9160 displays images and text. Display may be, for example, an LCD display, but is not limited thereto.
[0121] The memory 9140 can be a solid-state memory, such as a read-only memory (ROM), random access memory (RAM), a SIM card, etc. It can also be a memory that retains information even when power is off, can be selectively erased, and contains more data; examples of this type of memory are sometimes referred to as EPROMs. The memory 9140 can also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 may include an application / function storage unit 9142 for storing application programs and function programs or processes for executing the operation of the electronic device 9600 via the central processing unit 9100.
[0122] The memory 9140 may also include a data storage unit 9143 for storing data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various drivers for the electronic device for communication functions and / or for performing other functions of the electronic device (such as messaging applications, address book applications, etc.).
[0123] The communication module 9110 is a transmitter / receiver that sends and receives signals via the antenna 9111. The communication module 9110 (transmitter / receiver) is coupled to the central processing unit 9100 to provide input signals and receive output signals, which is the same as in a conventional mobile communication terminal.
[0124] Based on different communication technologies, multiple communication modules 9110 can be configured in the same electronic device, such as cellular network modules, Bluetooth modules, and / or wireless LAN modules. The communication module 9110 (transmitter / receiver) is also coupled to a speaker 9131 and a microphone 9132 via an audio processor 9130 to provide audio output via the speaker 9131 and receive audio input from the microphone 9132, thereby realizing typical telecommunications functions. The audio processor 9130 may include any suitable buffer, decoder, amplifier, etc. Additionally, the audio processor 9130 is coupled to a central processing unit 9100, enabling on-device recording via the microphone 9132 and on-device playback of stored audio via the speaker 9131.
[0125] Embodiments of this application also provide a computer-readable storage medium capable of implementing all steps of the closed-loop control method for an emotion-sensing robot with a server or client as the execution subject in the above embodiments. The computer-readable storage medium stores a computer program that, when executed by a processor, implements all steps of the closed-loop control method for the emotion-sensing robot with a server or client as the execution subject in the above embodiments. For example, when the processor executes the computer program, it implements the following steps: Step S101: Collect visual data stream, voice data stream, and tactile data stream input from the robot's perception system to generate multimodal sensor data. Preprocess the multimodal sensor data to obtain a standardized feature set. Perform deep feature learning on the standardized feature set based on an emotion feature extraction network to obtain an emotion state vector. Perform similarity matching between the emotion state vector and a preset emotion knowledge base to obtain scene description data. Perform context association analysis based on the scene description data to generate an interaction state vector. Step S102: Input the interaction state vector into the behavior decision network for strategy planning to obtain behavior control parameters, perform risk assessment on the behavior control parameters based on preset safety constraint rules to generate safety control indicators, construct an action planning strategy based on the safety control indicators to obtain an execution sequence graph, perform time sequence alignment processing on the execution sequence graph to generate a control instruction set, and input the control instruction set into the execution optimizer for resource scheduling to obtain a task scheduling matrix; Step S103: Construct a robot motion controller based on the task scheduling matrix to generate motion control parameters, perform real-time adjustment of the robot actuator based on the motion control parameters to obtain execution status data, input the execution status data into the memory management unit for experience learning to obtain strategy optimization parameters, and generate control update instructions based on the strategy optimization parameters to perform closed-loop optimization of the robot control system.
[0126] As described above, the computer-readable storage medium provided in this application embodiment achieves accurate emotion recognition through multimodal features and deep learning. A decision-making mechanism is constructed, combining safety constraints and resource scheduling to establish a reliable execution strategy. Closed-loop optimization is introduced, ensuring continuous improvement of control through experience learning and policy updates. This method effectively addresses the shortcomings of traditional technologies in emotion recognition, behavior control, and policy optimization, providing technical support for emotion-sensing robots.
[0127] Embodiments of this application also provide a computer program product capable of implementing all steps in the closed-loop control method for an emotion-sensing robot, where the execution subject is a server or client, as described in the above embodiments. When executed by a processor, this computer program / instruction implements the steps of the closed-loop control method for the emotion-sensing robot. For example, the computer program / instruction implements the following steps: Step S101: Collect visual data stream, voice data stream, and tactile data stream input from the robot's perception system to generate multimodal sensor data. Preprocess the multimodal sensor data to obtain a standardized feature set. Perform deep feature learning on the standardized feature set based on an emotion feature extraction network to obtain an emotion state vector. Perform similarity matching between the emotion state vector and a preset emotion knowledge base to obtain scene description data. Perform context association analysis based on the scene description data to generate an interaction state vector. Step S102: Input the interaction state vector into the behavior decision network for strategy planning to obtain behavior control parameters, perform risk assessment on the behavior control parameters based on preset safety constraint rules to generate safety control indicators, construct an action planning strategy based on the safety control indicators to obtain an execution sequence graph, perform time sequence alignment processing on the execution sequence graph to generate a control instruction set, and input the control instruction set into the execution optimizer for resource scheduling to obtain a task scheduling matrix; Step S103: Construct a robot motion controller based on the task scheduling matrix to generate motion control parameters, perform real-time adjustment of the robot actuator based on the motion control parameters to obtain execution status data, input the execution status data into the memory management unit for experience learning to obtain strategy optimization parameters, and generate control update instructions based on the strategy optimization parameters to perform closed-loop optimization of the robot control system.
[0128] As described above, the computer program product provided in this application achieves accurate emotion recognition through multimodal features and deep learning. A decision-making mechanism is constructed, combining safety constraints and resource scheduling to establish a reliable execution strategy. Closed-loop optimization is introduced, ensuring continuous improvement of control through experience learning and policy updates. This method effectively addresses the shortcomings of traditional technologies in emotion recognition, behavior control, and policy optimization, providing technical support for emotion-perceiving robots.
[0129] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0130] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (devices), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0131] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0132] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0133] Specific embodiments have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.
Claims
1. A closed-loop control method for an emotion-sensing robot, characterized in that, The method includes: The robot's perception system collects visual data streams, voice data streams, and tactile data streams to generate multimodal sensor data. The multimodal sensor data is preprocessed to obtain a standardized feature set. Based on an emotion feature extraction network, the standardized feature set is subjected to deep feature learning to obtain an emotion state vector. The emotion state vector is matched with a preset emotion knowledge base to obtain scene description data. The scene description data is used to perform context association analysis to generate an interaction state vector. The interaction state vector is input into the behavior decision network for strategy planning to obtain behavior control parameters. Based on preset safety constraint rules, the behavior control parameters are risk-assessed to generate safety control indicators. Based on the safety control indicators, an action planning strategy is constructed to obtain an execution sequence graph. The execution sequence graph is time-aligned to generate a control instruction set. The control instruction set is input into the execution optimizer for resource scheduling to obtain a task scheduling matrix. Based on the task scheduling matrix, a robot motion controller is constructed to generate motion control parameters. Based on the motion control parameters, the robot actuator is adjusted in real time to obtain execution status data. The execution status data is input into the memory management unit for experience learning to obtain strategy optimization parameters. Based on the strategy optimization parameters, control update instructions are generated to perform closed-loop optimization of the robot control system.
2. The closed-loop control method for the emotion-sensing robot according to claim 1, characterized in that, The visual data stream, voice data stream, and tactile data stream input by the robot's perception system generate multimodal sensor data. The multimodal sensor data is preprocessed to obtain a standardized feature set, including: The image data stream acquired by the robot vision sensor is spatially filtered to obtain a filtered image set. The acoustic signal stream acquired by the voice sensor is frequency transformed to obtain an acoustic feature spectrum. The pressure signal stream acquired by the tactile sensor is sampled and quantized to obtain a tactile feature sequence. Based on a multimodal synchronization mechanism, the filtered image set, the acoustic feature spectrum, and the tactile feature sequence are timestamped to generate multimodal sensor data. The multimodal sensor data is input into a feature preprocessing network for data normalization to obtain a normalized feature set. The normalized feature set is then subjected to noise suppression and dimensionality transformation to obtain a feature dimensionality reduction matrix. Based on a feature selection algorithm, the feature dimensionality reduction matrix is used to filter features and generate a standardized feature set.
3. The closed-loop control method for the emotion-sensing robot according to claim 1, characterized in that, The process involves performing deep feature learning on the standardized feature set using an emotion feature extraction network to obtain an emotion state vector, matching the emotion state vector with a preset emotion knowledge base to obtain scene description data, and generating an interaction state vector based on contextual association analysis of the scene description data. This includes: A standardized feature set is input into an emotion feature extraction network for hierarchical feature mapping to obtain a hierarchical feature group. An attention mechanism is applied to the hierarchical feature group to generate a weight distribution map. The hierarchical feature group is then weighted and fused based on the weight distribution map to obtain an emotion state vector. The emotion state vector is then compared with labeled samples in a preset emotion knowledge base to generate a matching score matrix. A scene semantic parser is constructed based on the matching score matrix to obtain a semantic mapping table. Scene elements are semantically labeled based on the semantic mapping table to generate scene description data. The scene description data is input into a context analysis model to perform temporal association modeling to obtain an interaction state vector.
4. The closed-loop control method for the emotion-sensing robot according to claim 1, characterized in that, The step of inputting the interaction state vector into the behavior decision network for policy planning to obtain behavior control parameters, and then performing risk assessment on the behavior control parameters based on preset safety constraint rules to generate safety control indicators includes: The interaction state vector is analyzed in multiple layers through a behavior decision network to obtain a decision feature map. The decision feature map is prioritized to generate a task sequence list. Based on a preset behavior template, the task sequence list is matched with a strategy to obtain a candidate strategy set. The candidate strategy set is then subjected to conflict detection to generate behavior control parameters. The behavior control parameters are subjected to safety boundary checks to obtain a set of constraint parameters. Based on preset safety constraint rules, the constraint parameter set is subjected to risk measurement to generate a risk score table. The risk score table is then filtered through a safety threshold to obtain safety control indicators.
5. The closed-loop control method for the emotion-sensing robot according to claim 1, characterized in that, The process of constructing an action planning strategy based on the safety control indicators to obtain an execution sequence graph, performing time-sequence alignment processing on the execution sequence graph to generate a control instruction set, and inputting the control instruction set into the execution optimizer for resource scheduling to obtain a task scheduling matrix includes: The safety control indicators are input into the action planner to decompose the action into a primitive action set. The primitive action set is combined and reconstructed based on the preset action syntax rules to generate an action chain sequence. The action chain sequence is then passed through the spatiotemporal constraint checker to resolve conflicts and obtain an execution sequence graph. The execution sequence graph is then time-marked to generate a time sequence relationship table. Based on the timing relationship table, a synchronization controller is constructed to obtain a control timing diagram. The control timing diagram is then decomposed into a task to generate a control instruction set. The control instruction set is then input into the execution optimizer to allocate computational resources and obtain a task scheduling matrix.
6. The closed-loop control method for the emotion-sensing robot according to claim 1, characterized in that, The step of constructing a robot motion controller based on the task scheduling matrix to generate motion control parameters, and then performing real-time adjustment of the robot actuator based on the motion control parameters to obtain execution state data includes: The task scheduling matrix is input into the trajectory planner to generate a path and obtain a motion trajectory map. The motion trajectory map is then subjected to dynamic constraint verification to generate a constraint parameter set. Based on the constraint parameter set, a motion controller is constructed to obtain motion control parameters. The motion control parameters are then mapped hierarchically according to execution priority to generate an execution instruction table. The robot joint actuators are controlled in real time according to the execution instruction table to obtain joint state data. The joint state data is then monitored through a sensor feedback network to generate execution state data. The execution state data is then verified in real time to generate a state feedback signal.
7. The closed-loop control method for the emotion-sensing robot according to claim 1, characterized in that, The step of inputting the execution state data into the memory management unit for experience learning to obtain strategy optimization parameters, and generating control update instructions based on the strategy optimization parameters to perform closed-loop optimization of the robot control system includes: The execution state data is encoded using a memory encoder to obtain a state feature group. The state feature group is then subjected to empirical pattern extraction to generate an empirical sample set. Based on a preset learning rule, reinforcement learning is performed on the empirical sample set to obtain policy optimization parameters. The policy optimization parameters are then input into an evaluation model to perform performance evaluation and generate an optimization index table. Based on the optimization index table, a feedback compensator is constructed to obtain a set of compensation parameters. The set of compensation parameters is adaptively adjusted to generate control update instructions. The control update instructions are then passed through a closed-loop controller to correct the parameters and obtain a control correction amount. Based on the control correction amount, the robot control system is optimized in a closed loop.
8. A closed-loop control device for an emotion-sensing robot, characterized in that, The device includes: The interaction analysis module is used to collect visual data streams, voice data streams, and tactile data streams input from the robot's perception system to generate multimodal sensor data. The multimodal sensor data is preprocessed to obtain a standardized feature set. Based on an emotion feature extraction network, deep feature learning is performed on the standardized feature set to obtain an emotion state vector. The emotion state vector is matched with a preset emotion knowledge base to obtain scene description data. Context association analysis is performed based on the scene description data to generate an interaction state vector. The task control module is used to input the interaction state vector into the behavior decision network for strategy planning to obtain behavior control parameters, perform risk assessment on the behavior control parameters based on preset safety constraint rules to generate safety control indicators, construct action planning strategies based on the safety control indicators to obtain an execution sequence graph, perform time sequence alignment processing on the execution sequence graph to generate a control instruction set, and input the control instruction set into the execution optimizer for resource scheduling to obtain a task scheduling matrix. The strategy formulation module is used to construct a robot motion controller based on the task scheduling matrix to generate motion control parameters, to adjust the robot actuator in real time based on the motion control parameters to obtain execution status data, to input the execution status data into the memory management unit for experience learning to obtain strategy optimization parameters, and to generate control update instructions based on the strategy optimization parameters to perform closed-loop optimization of the robot control system.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the closed-loop control method for the emotion-sensing robot according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the closed-loop control method for the emotion-sensing robot according to any one of claims 1 to 7.