A scene switching method and device of a robot and the robot

CN122500677APending Publication Date: 2026-08-04BEIJING KEYI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING KEYI TECH CO LTD
Filing Date
2026-04-01
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

具体地,当前机器人的场景切换方式,是把复杂任务简化为一套非此即彼的硬性触发指令,跳过了生物智能所必需的、基于持续感知和内部状态演变的、渐进式的“思考”与“准备”过程,从而导致行为生硬、场景切换突兀,缺乏连贯性,给用户的体验较差

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122500677A_ABST
    Figure CN122500677A_ABST
Patent Text Reader

Abstract

This application provides a method, apparatus, and robot for scene switching, relating to the field of robot technology. The method includes: extracting features from multimodal data of a target in different frames to obtain frame feature information, and determining the target attention value of the target in the current frame based on the frame feature information; determining the stimulus event of the current frame based on the multimodal data, and determining the emotional state information of the current frame based on the weight of the stimulus event; determining the target scene tendency value of the current frame based on the scene tendency value, target attention value, and emotional state information of the previous frame; determining the target scene according to the target scene tendency value, comparing the target scene with the current scene, and determining whether to perform scene switching based on a first comparison result. This application can improve the coherence of scene switching in robots, making scene switching more natural and improving user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robotics, and in particular to a method, apparatus, and robot for scene switching. Background Technology

[0002] With the continuous development of technology, various service robots have gradually entered thousands of households and have been widely promoted and applied.

[0003] Currently, to meet diverse user needs, service robots are typically deployed in various scenarios, such as office work and interactive environments. The decision-making layer of service robots usually uses discrete rules based on current multimodal data to determine scenario switching. Specifically, the current method of scenario switching for robots simplifies complex tasks into a set of hard-triggered commands that are either one way or the other, skipping the gradual "thinking" and "preparation" process necessary for biological intelligence—based on continuous perception and internal state evolution. This results in stiff behavior, abrupt scenario transitions, a lack of coherence, and a poor user experience. Summary of the Invention

[0004] This application provides a method, apparatus, and robot for scene switching, which improves the continuity of scene switching, makes scene switching more natural, and enhances the user experience.

[0005] On one hand, embodiments of this application provide a scene switching method for a robot, including: Feature extraction is performed on the multimodal data of the target in different frames to obtain frame feature information, and the target attention value of the target in the current frame is determined based on the frame feature information; The stimulus event of the current frame is determined based on the multimodal data, and the emotional state information of the current frame is determined based on the weight of the stimulus event. The stimulus event is an event that causes a change in the emotional state. The target scene tendency value of the current frame is determined based on the scene tendency value of the previous frame, the target attention value, and the emotional state information. The target scene is determined based on the target scene tendency value, and the target scene is compared with the current scene. Based on the first comparison result, it is determined whether to switch scenes.

[0006] In this embodiment, feature extraction is performed on the multimodal data of the target in different frames to obtain frame feature information. Based on the frame feature information, the target attention value of the target in the current frame is determined, thereby improving the accuracy and robustness of the target attention value calculation. This application determines the stimulus event of the current frame based on multimodal data and determines the emotional state information of the current frame based on the weight of the stimulus event, thereby improving the accuracy of the emotional state information of the current frame. This application determines the target scene tendency value of the current frame based on the scene tendency value, target attention value, and emotional state information of the previous frame, thereby reflecting the scene bias of the current frame. Furthermore, the target scene is determined according to the target scene tendency value, and the target scene is compared with the current scene. Based on the first comparison result, it is determined whether to switch scenes, thereby improving the coherence of the robot's scene switching, making scene switching more natural and possessing "intentional" characteristics, and improving the user experience.

[0007] Optionally, determining the target attention value of the target in the current frame based on the frame feature information includes: Based on the frame feature information, the attention value of the target in different frames is determined; The target attention value is determined based on the target's attention values ​​in different frames and the total number of frames.

[0008] In this embodiment, the attention value of the target in different frames is determined based on frame feature information, and the target attention value is determined based on the target attention value in different frames and the total number of frames, so as to improve the accuracy of the target attention value.

[0009] Optionally, determining the stimulus event of the current frame based on the multimodal data includes some or all of the following steps: Based on the frame feature information, the burst level information of the current frame is determined. If the burst level information of the current frame exceeds the set level threshold, the stimulus event is generated. Based on the attention values ​​of the target in different frames, the duration of the target as the primary attention target is determined. If the duration of the target as the primary attention target exceeds a first duration threshold, the stimulus event is generated. Emotion recognition is performed on the speech data of multimodal data from different frames. Based on the emotion recognition results, the duration of the speech data corresponding to different emotions is determined. If there is an emotion that exceeds a second duration threshold, the stimulus event is generated. Target detection is performed on image data of multimodal data from different frames. Based on the target detection results, a first duration corresponding to image data that does not include the target is determined. If the first duration exceeds a third duration threshold, the stimulus event is generated. Based on the pressure data from multimodal data of different frames, the duration of touch on different parts of the robot is determined. If the duration exceeds the fourth duration threshold, the stimulus event is generated.

[0010] Optionally, before determining the target scene tendency value of the current frame, the method further includes: Based on the frame feature information, the target burst level information is determined; Determining the target scene tendency value for the current frame based on the scene tendency value of the previous frame, the target attention value, and the emotional state information includes: The target scene tendency value is determined based on the scene tendency value of the previous frame of the current frame, the target attention value, the target suddenness level information, and the emotional state information.

[0011] In this embodiment, target burst level information is determined based on frame feature information, thereby achieving millisecond-level response to sudden stimuli and maintaining stability through a smoothing mechanism. Furthermore, this application determines the target scene tendency value based on the scene tendency value, target attention value, target burst level information, and emotional state information of the previous frame, thus more accurately reflecting the scene bias of the current frame. This improves the accuracy and coherence of scene switching, making scene switching more natural and possessing "intentional" characteristics, thereby enhancing the user experience.

[0012] Optionally, determining the target burst level information based on the frame feature information includes: Based on the frame feature information and the second duration, the feature change rate of the target in different frames is determined, where the second duration is the duration between two adjacent frames; The burst level information of the target is determined based on the characteristic change rate of the target in different frames.

[0013] In this embodiment, based on frame feature information and a second duration, the feature change rate of the target in different frames is determined to reflect the changes in frame feature information within the second duration. Furthermore, this application determines the target burst level information based on the target's feature change rate in different frames, thereby improving the accuracy of the target burst level information.

[0014] Optionally, determining the emotional state information of the current frame based on the weight of the stimulus event includes: If the current time is within the task's execution time period, then determine the initial emotional state information and the termination emotional state information for the execution time period. Based on the duration corresponding to the execution time period, the third duration, the initial emotional state information, and the termination emotional state information, the first emotional state information is determined, and the third duration is determined based on the current time and the earliest time of the execution time period; Based on the weights of the stimulus events, the second emotional state information is determined; The emotional state information is determined based on the first emotional state information and the second emotional state information.

[0015] In this embodiment, when the current time is within the task's execution period, a first emotional state is determined based on the duration of the task's execution period, a third duration, the initial emotional state information of the execution period, and the termination emotional state information. This first emotional state reflects the emotional state of the current frame during the transition from the initial emotional state to the termination emotional state within the execution period. Furthermore, this application determines a second emotional state based on the weight of the stimulus event, reflecting the degree of influence of the stimulus event triggered in the current frame on the emotional state. This application determines the emotional state information of the current frame based on the first and second emotional state information, improving the accuracy of the emotional state information of the current frame.

[0016] Optionally, the first emotional state information can be determined using the following formula. : ; in, This is information about the initial emotional state. To terminate emotional state information, This refers to the earliest time of the task's execution period. This refers to the latest time during which the task will be executed. For the current time, .

[0017] Optionally, determining the emotional state information of the current frame based on the weight of the stimulus event includes: If the current time is not within the task's execution time period, the second emotional state information is determined based on the weight of the stimulus event. The second emotional state information is compared with a set threshold to obtain a second comparison result; If the second comparison result indicates that the second emotional state information is less than a set threshold, then it is determined that the robot is in an idle state, and based on the third emotional state information when the robot is first detected to be in an idle state, the set emotional state information, and the fourth duration, the fourth emotional state information is determined, wherein the fourth duration is determined based on the current time and the time when the robot is first detected to be in an idle state; The emotional state information is determined based on the second emotional state information and the fourth emotional state information.

[0018] In this embodiment, if the current time is not within the task's execution time period and the second emotional state information is less than a set threshold, the robot is determined to be in an idle state. Based on the third emotional state information when the robot is first detected to be in an idle state, the set emotional state information, and the fourth duration, a fourth emotional state information is determined to reflect the emotional state of the current frame during the transition from the third emotional state to the set emotional state within the idle time period. Furthermore, this application determines the emotional state information of the current frame based on the second and fourth emotional state information, improving the accuracy of the emotional state information of the current frame.

[0019] Optionally, after comparing the second emotional state information with a set threshold to obtain a second comparison result, the method further includes: If the second comparison result indicates that the second emotional state information is greater than or equal to a set threshold, then the second emotional state information is used as the emotional state information of the current frame.

[0020] Optionally, the frame feature information can be determined by the following method: Based on the scene corresponding to different frames, the target and the first correspondence relationship, the baseline feature information of the frame feature information is determined, and the first correspondence relationship is the correspondence relationship between each scene, each target and each baseline feature information. Static feature extraction is performed on the image data of multimodal data from different frames to determine the static feature information of the frame feature information; Dynamic feature extraction is performed on the image data of multimodal data from different frames to determine the dynamic feature information of the frame feature information.

[0021] In this embodiment of the application, based on the scene, target and first correspondence corresponding to different frames, the baseline feature information of the frame feature information is determined, and static feature extraction and dynamic feature extraction are performed on the image data of multimodal data of different frames respectively to determine the static feature information and dynamic feature information of the frame feature information, thereby improving the richness of the frame feature information.

[0022] Optionally, the step of extracting static features from the image data of multimodal data of different frames to determine the static feature information of the frame feature information includes some or all of the following steps: Target detection is performed on image data from different frames to determine the target bounding boxes of the target in different frames. Based on the center point coordinates of the target bounding boxes in different frames and the center point coordinates of the image data in different frames, the centrality feature information of the static feature information is determined. Based on image data from different frames, a deep learning method is used to determine the target angle of the target in different frames, and based on the target angle of the target in different frames, the orientation feature information of the static feature information is determined, wherein the target angle is the angle between the target normal and the camera's line of sight; Target detection and key point detection are performed on image data of different frames to determine the key points of the target in different frames. Based on the key points of the target in different frames and the depth image data of multimodal data in different frames, the distance feature information of the static feature information is determined. Target detection is performed on image data from different frames to determine the number of times the target is detected, and confidence feature information of the static feature information is determined based on the number of times.

[0023] Optionally, the dynamic feature extraction of image data from multimodal data of different frames to determine the dynamic feature information of the frame feature information includes some or all of the following steps: Target detection is performed on image data from different frames to determine the center point coordinates of the target bounding box in different frames. Based on the center point coordinates of the target bounding box in different frames and the second duration, the velocity feature information of the dynamic feature information is determined. Based on the velocity feature information and the second duration, the acceleration feature information of the dynamic feature information is determined; The variance of the velocity feature information is calculated to determine the rhythmic feature information of the dynamic feature information.

[0024] Optionally, determining whether to switch scenes based on the first comparison result includes: If the first comparison result indicates that the target scene and the current scene are the same, then the current scene remains unchanged; If the first comparison result indicates that the target scene and the current scene are different, then the current scene is switched to the target scene.

[0025] Optionally, after determining whether to switch scenes based on the first comparison result, the method further includes: Based on the attention value of the target in the current frame, determine the main attention target of the current frame; Based on the target scene, the emotional state information, the main attention target of the current frame, and the historical task information, a target task chain is selected from multiple preset task chains and executed.

[0026] On one hand, embodiments of this application provide a scene switching device for a robot, including: The first determining module is used to extract features from the multimodal data of the target in different frames to obtain frame feature information, and to determine the target attention value of the target in the current frame based on the frame feature information. The second determining module is used to determine the stimulus event of the current frame based on the multimodal data, and to determine the emotional state information of the current frame based on the weight of the stimulus event, wherein the stimulus event is an event that causes a change in the emotional state. The third determining module is used to determine the target scene tendency value of the current frame based on the scene tendency value of the previous frame of the current frame, the target attention value, and the emotional state information. The switching module is used to determine the target scene based on the target scene tendency value, compare the target scene with the current scene, and determine whether to switch scenes based on the first comparison result.

[0027] On one hand, embodiments of this application provide a robot, including at least one processor and at least one memory, wherein the memory stores a computer program, and when the program is executed by the processor, the processor performs the steps of the scene switching method of the robot described above.

[0028] On one hand, embodiments of this application provide a computer-readable storage medium storing a computer program executable by a computer device, which, when run on the computer device, causes the computer device to perform the steps of the scene switching method for the robot described above.

[0029] On one hand, embodiments of this application provide a computer program product, the computer program product including a computer program stored on a computer-readable storage medium, the computer program including program instructions, which, when executed by a computer device, cause the computer device to perform the steps of the above-described robot scene switching method. Attached Figure Description

[0030] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0031] Figure 1 A flowchart illustrating a scene switching method for a robot provided in an embodiment of this application; Figure 2 A detailed flowchart of a scene switching method for a robot provided in an embodiment of this application; Figure 3A flowchart illustrating a method for determining frame feature information provided in an embodiment of this application; Figure 4 A flowchart illustrating a method for determining the target attention value of a target in the current frame, provided in an embodiment of this application; Figure 5 A flowchart illustrating a method for determining the target attention value of any target in the current frame, provided in an embodiment of this application; Figure 6 A flowchart illustrating a method for determining the burst level information of the current frame, provided in an embodiment of this application; Figure 7 A flowchart illustrating a method for determining the emotional state information of the current frame, provided in an embodiment of this application; Figure 8 A flowchart illustrating another method for determining the emotional state information of the current frame, provided in an embodiment of this application; Figure 9 A detailed flowchart illustrating a method for determining the emotional state information of the current frame, provided in an embodiment of this application; Figure 10 A flowchart illustrating a method for determining the outbreak level information of a target, provided in an embodiment of this application; Figure 11 A flowchart illustrating a method for determining burst level information of any frame, provided in an embodiment of this application; Figure 12 A detailed flowchart illustrating another scene switching method for a robot provided in this application embodiment; Figure 13 A schematic diagram of the structure of a scene switching device for a robot provided in an embodiment of this application; Figure 14 This is a schematic diagram of the structure of a robot provided in an embodiment of this application. Detailed Implementation

[0032] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this application. Obviously, the described embodiments are only some embodiments of the technical solutions of this application, and not all embodiments. Based on the embodiments recorded in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the technical solutions of this application.

[0033] The data collection, dissemination, and use in this application all comply with relevant national laws and regulations.

[0034] Before introducing the scene switching method for a robot provided in the embodiments of this application, the technical background of the embodiments of this application will be described in detail below for ease of understanding.

[0035] With the continuous development of technology, various service robots have gradually entered thousands of households and have been widely promoted and applied.

[0036] Service robots in current home and office environments need to simultaneously manage multimodal perception, attention management, behavioral responses, and emotional consistency during interactions. Traditional solutions typically separate visual attention, emotion assessment, and task decision-making, leading to the following problems: Problem 1: The separation between the perception and object layers makes attention difficult to stabilize.

[0037] Existing systems often rely on single or limited visual / depth sensors and lack sufficient fusion of different perceptual results. Target tracking using only RGB (Red / Green / Blue) images is prone to jitter under occlusion, lighting changes, or rapid movement; depth information alone is insufficient to obtain the semantic attributes of the target, as it only contains geometric shape and position data, lacking visual features such as color, texture, and pattern to determine "what the object is," thus making it difficult to independently identify semantic attributes. The heterogeneity of multi-source data in terms of time sequence and format further affects the accuracy and robustness of attention computation.

[0038] Question 2: Attention allocation lacks multidimensionality and historical consistency.

[0039] Existing attention mechanisms mostly employ fixed rules or single-factor driven approaches, making it difficult to simultaneously integrate spatial saliency, dynamic behavioral characteristics, and historical attentional trends. In multi-objective competitive scenarios, this can easily lead to frequent attention switching and focus shifts, disrupting the continuity of interaction. Among these, fixed-rule mechanisms rely solely on immediate reflexes based on a single condition.

[0040] Question 3: The emotional system lacks a closed-loop mechanism.

[0041] Most robot emotion modules are only used for expression rendering (facial expression / voice style), and do not form an effective feedback loop with attention and decision-making. Emotions are difficult to become an "internal state" and cannot help with scene judgment and task rhythm control.

[0042] Question 4: The decision-making process lacks biomimetic continuous judgment and multimodal consistent output.

[0043] Current robot decision-making methods simplify complex tasks into a set of hard-and-white trigger commands, skipping the gradual "thinking" and "preparation" processes necessary for biological intelligence—based on continuous perception and internal state evolution. They lack continuous judgment based on the accumulation of attentional evidence and emotional trends, resulting in abrupt and inconsistent scene transitions and a poor user experience. Furthermore, actions, speech, and facial expressions are triggered independently by their respective modules, lacking a unified control signal, leading to asynchronous and unnatural output.

[0044] Therefore, there is an urgent need for a hierarchical, collaborative, and real-time closed-loop decision-making architecture, namely a scene switching method for robots, which integrates object attention, emotion evolution, and scene / task decision-making, enabling robots to possess characteristics similar to humans such as "rapid response - emotion regulation - rational decision-making".

[0045] To address the aforementioned issues, this application proposes a method, apparatus, and robot for scene switching, which aims to improve the continuity of scene switching, make scene switching more natural, and enhance the user experience.

[0046] The following is for reference. Figure 1 The flowchart illustrating a scene switching method for a robot is shown below, which explains the technical solution provided in the embodiments of this application: Step 101: Extract features from the multimodal data of the target in different frames to obtain frame feature information, and determine the target attention value of the target in the current frame based on the frame feature information.

[0047] The target can be one or more, including targets such as Face, Hand, Body, and Object. If there are multiple targets, the target attention value of the target in the current frame includes the target attention value of each target in the current frame. If there is only one target, the target attention value of the target in the current frame includes the target attention value of that target in the current frame. For example, if the targets include Face and Hand, the target attention value of the target in the current frame includes the target attention values ​​of Face and Hand in the current frame. If the target includes Object, the target attention value of the target in the current frame includes the target attention value of Object in the current frame.

[0048] In this embodiment of the application, before obtaining the frame feature information, the method further includes: acquiring multimodal data of the target over a set time period from different collectors of the robot, wherein the multimodal data of the target over the set time period includes multimodal data of the target in different frames.

[0049] The time period is determined based on historical time and the current time, and the duration of the set time period (the duration of the historical time and the current time) can be set according to the actual situation. For example, if the current time is t, the duration of the set time period is... Then set the time period as The robot's data acquisition devices include pressure sensors, cameras, and voice acquisition units. Multimodal data includes image data, voice data, pressure data, and other data.

[0050] In this embodiment, the cerebellar attention module in the robot is used to extract features from the multimodal data of the target in different frames, obtain frame feature information, and determine the target attention value of the target in the current frame based on the frame feature information.

[0051] Step 102: Determine the stimulus event of the current frame based on multimodal data, and determine the emotional state information of the current frame based on the weight of the stimulus event.

[0052] In this embodiment of the application, the emotion model module (Valence-Arousal-Dominance, VAD) in the robot is used to determine the stimulus event of the current frame based on multimodal data, and to determine the emotional state information of the current frame based on the weight of the stimulus event.

[0053] Among them, the stimulus event is an event that causes a change in emotional state. There can be one or more stimulus events, and each stimulus event corresponds to a weight.

[0054] In this embodiment, the emotional state information of frame t is represented by the following formula. : ;Formula (1) in, For frame t, the level of pleasure / positive / negative tendency. The arousal level / activation level of frame t. The degree of dominance / dominance of frame t. , and The value of is generally normalized to the interval [-1, 1].

[0055] Step 103: Determine the target scene tendency value of the current frame based on the scene tendency value, target attention value, and emotional state information of the previous frame of the current frame.

[0056] In this embodiment of the application, the target scene tendency value of the current frame is determined based on the scene tendency value, target attention value, and emotional state information of the previous frame of the current frame. This includes: determining the target scene tendency value of the current frame based on the scene tendency value, target attention value, emotional state information, and their respective weights of the previous frame of the current frame.

[0057] Step 104: Determine the target scene based on the target scene tendency value, compare the target scene with the current scene, and determine whether to switch scenes based on the first comparison result.

[0058] In this embodiment of the application, the brain decision-making module in the robot is used to determine the target scene tendency value of the current frame based on the scene tendency value, target attention value and emotional state information of the previous frame of the current frame, determine the target scene based on the target scene tendency value, compare the target scene with the current scene, and determine whether to switch scenes based on the first comparison result.

[0059] Optionally, determining the target scenario based on the target scenario tendency value includes: taking the scenario corresponding to the set tendency value range to which the target scenario tendency value belongs as the target scenario.

[0060] Optionally, determining whether to switch scenes based on the first comparison result includes: If the first comparison result indicates that the target scene and the current scene are the same, then the current scene remains unchanged; If the first comparison result indicates that the target scene and the current scene are different, then the current scene will be switched to the target scene.

[0061] In this embodiment, feature extraction is performed on the multimodal data of the target in different frames to obtain frame feature information. Frame feature information includes baseline feature information, static feature information, and dynamic feature information. Based on the frame feature information, the target attention value of the target in the current frame is determined, thereby improving the accuracy and robustness of the target attention value calculation. This application determines the stimulus event of the current frame based on multimodal data and determines the emotional state information of the current frame based on the weight of the stimulus event, thereby improving the accuracy of the emotional state information of the current frame. Based on the scene tendency value, target attention value, and emotional state information of the previous frame, this application determines the target scene tendency value of the current frame, thereby reflecting the scene bias of the current frame. Furthermore, the target scene is determined according to the target scene tendency value, and the target scene is compared with the current scene. Based on the first comparison result, it is determined whether to switch scenes, thereby improving the coherence of the robot's scene switching, making scene switching more natural and possessing "intentional" characteristics, and improving the user experience.

[0062] The following will provide a detailed explanation of the specific steps involved in the scene switching method for a robot provided above, such as... Figure 2 As shown: Step 201: Extract features from the multimodal data of the target in different frames to obtain frame feature information.

[0063] Step 202: Determine the target attention value of the target in the current frame based on frame feature information.

[0064] Step 203: Determine the stimulus event of the current frame based on multimodal data.

[0065] Step 204: Determine the emotional state information of the current frame based on the weight of the stimulus event.

[0066] Step 205: Determine the target scene tendency value of the current frame based on the scene tendency value, target attention value, and emotional state information of the previous frame of the current frame.

[0067] Step 206: Select the scene corresponding to the set tendency value range to which the target scene tendency value belongs as the target scene.

[0068] Step 207: Determine whether the target scene and the current scene are the same. If they are the same, proceed to step 208; otherwise, proceed to step 209.

[0069] Step 208: Keep the current scene unchanged.

[0070] Step 209: Switch the current scene to the target scene.

[0071] In this embodiment, feature extraction is performed on the multimodal data of the target in different frames to obtain frame feature information. The frame feature information includes the frame feature information of each target in each frame. Figure 3 A flowchart illustrating a method for determining frame feature information provided in an embodiment of this application is shown below. Figure 3 As shown, step 201 above includes at least the following steps 301-303: Step 301: Determine the baseline feature information of the frame feature information based on the scene, target and first correspondence corresponding to different frames.

[0072] The first correspondence is the correspondence between each scene, each target, and each baseline feature information. The baseline feature information of the frame feature information includes the baseline feature information of each target in each frame.

[0073] In this embodiment of the application, the baseline feature information of any target in any frame is determined by the following method: based on the scene corresponding to any frame, any target and the first correspondence, the baseline feature information of any target in any frame is determined.

[0074] The baseline feature information of any target in any frame signifies the "default importance" of that target within the corresponding scene of that frame. In an "office" scene, the baseline feature information is higher when the target is "Face"; in a "play / interaction" scene, the baseline feature information can be appropriately increased when the target is "Hand" or "Object". The scene template parameters are issued by the robot's brain decision-making module. Decide.

[0075] In this embodiment, the robot's scene set S = {office, interactive, standby...}, and the scene corresponding to frame t can be denoted as... The robot's target set Let the subscript c represent a target.

[0076] Step 302: Extract static features from the image data of multimodal data of different frames to determine the static feature information of frame feature information.

[0077] The static feature information of the frame feature information includes the static feature information of each target in each frame. The static feature information includes some or all of the following: centrality feature information, orientation feature information, distance feature information, and confidence feature information.

[0078] Step 303: Dynamic feature extraction is performed on the image data of multimodal data from different frames to determine the dynamic feature information of the frame feature information.

[0079] The dynamic feature information of the frame feature information includes the dynamic feature information of each target in each frame. The dynamic feature information includes some or all of the following: velocity feature information, acceleration feature information, and rhythm feature information.

[0080] In this embodiment, the centrality information in the static feature information is determined as follows: target detection is performed on image data of different frames to determine the target bounding boxes of the target in different frames, and the centrality information of the static feature information is determined based on the center point coordinates of the target bounding boxes in different frames and the center point coordinates of the image data in different frames. The centrality information of the static feature information includes the feature information of each target in each frame.

[0081] Specifically, the centrality feature information of any target in any frame is determined by the following method: target detection is performed on the image data of any frame to determine the target bounding box of any target in any frame, and the centrality feature information of any target in any frame is determined based on the center point coordinates of the target bounding box of any target in any frame and the center point coordinates of the image data of any frame.

[0082] The center point coordinates of the image data refer to the fixed center point coordinates of the entire display screen, which are determined by the image resolution. For example, the center point of a 640×480 screen is (320, 240).

[0083] Optionally, object detection is performed on any frame of image data to determine the bounding box of any object in any frame. This includes: inputting the image data of any frame into an object detection model for object detection, and obtaining the bounding box of any object in any frame output by the object detection model. The object detection model can be pre-trained using machine learning methods (such as supervised learning methods). The base model used to train the object detection model can be various models with detection capabilities, such as convolutional neural network models, neural network models, etc.

[0084] Optionally, based on the center point coordinates of the target bounding box of any target in any frame and the center point coordinates of the image data of any frame, the centrality feature information of any target in any frame is determined, including: based on the center point coordinates of the target bounding box of any target in any frame and the center point coordinates of the image data of any frame, the Euclidean distance is determined, and the centrality feature information of any target in any frame is determined based on the Euclidean distance.

[0085] In this embodiment, after determining the Euclidean distance, the Euclidean distance is normalized to obtain a normalized distance. Based on this normalized distance, the centrality feature information of any target in any frame is determined. Specifically, the normalized distance is determined by dividing the Euclidean distance by half the length of the diagonal of the image data in any frame.

[0086] In this embodiment, the centrality feature information of target c in frame t is determined by the following formula. : ;Formula (2) in, The normalized distance in frame t is determined based on the center coordinates of the target bounding box of target c in frame t and the center coordinates of the image data in frame t, i.e., the normalized distance from the target center to the image center. .

[0087] In this embodiment of the application, the orientation feature information of static feature information is determined in the following way: based on image data of different frames, the target angle of the target in different frames is determined by deep learning method, and the orientation feature information of static feature information is determined based on the target angle of the target in different frames. The target angle is the angle between the target normal and the camera's line of sight. The orientation feature information of static feature information includes the orientation feature information of each target in each frame.

[0088] Specifically, the orientation feature information of any target in any frame is determined in the following way: based on the image data of any frame, the target angle of any target in any frame is determined using deep learning methods, and based on the target angle of any target in any frame, the orientation feature information of any target in any frame is determined.

[0089] The specific process of determining the angle between the target normal and the camera's line of sight using deep learning methods based on image data is existing technology and will not be described in detail here.

[0090] Optionally, based on the target angle of any target in any frame, the orientation feature information of any target in any frame is determined, including: based on the target angle of any target in any frame, the orientation feature information of any target in any frame is determined using a cosine function.

[0091] In this embodiment, the orientation feature information of target c in frame t is determined by the following formula. : ;Formula (3) in, Let the target angle of target c in frame t be the angle between the target normal of target c in frame t and the camera's line of sight (approximately when facing directly). ).

[0092] In this embodiment, the distance feature information of static feature information is determined as follows: target detection and keypoint detection are performed on image data of different frames to determine the keypoints of the target in different frames, and the distance feature information of static feature information is determined based on the keypoints of the target in different frames and the depth image data of multimodal data in different frames. The distance feature information of static feature information includes the distance feature information of each target in each frame.

[0093] In this embodiment, the distance feature information of any target in any frame is determined as follows: target detection and keypoint detection are performed on the image data of any frame to determine the keypoints of any target in any frame; based on the keypoints of any target in any frame and the depth image data of any frame in the multimodal data, the depth value of each keypoint is determined; based on the depth value of each keypoint and a set slope, the distance feature information of any target in any frame is determined. The image data of any frame is RGB image data. The image feature information can also be referred to as size feature information.

[0094] Optionally, object detection and keypoint detection are performed on any frame of image data to determine the keypoints of any object in any frame. This includes: inputting the image data of any frame into the detection model for object detection and keypoint detection, and obtaining the keypoints of any object in any frame output by the detection model. The detection model can be pre-trained using machine learning methods (such as supervised learning methods). The base model used to train the detection model can be various models with detection functions, such as convolutional neural network models, neural network models, etc.

[0095] Optionally, based on the depth values ​​of each key point and a set slope, the distance feature information of any target in any frame is determined, including: determining the target distance information based on the depth values ​​of each key point, and determining the distance feature information of any target in any frame based on the target distance information and a set slope.

[0096] In this embodiment, the distance feature information of target c in frame t is determined by the following formula. : ;Formula (4) in, The slope can be set according to the actual situation. The target distance information for target c in frame t is the "distance indicator" corresponding to the depth or frame size.

[0097] In this embodiment, the confidence feature information of static feature information is determined as follows: target detection is performed on image data of different frames to determine the number of times the target is detected, and the confidence feature information of static feature information is determined based on the number of times. The confidence feature information of static feature information includes the confidence feature information of each target in each frame.

[0098] In this embodiment of the application, the confidence feature information of any target in any frame is determined by the following method: target detection is performed on the image data of multiple frames before any frame in the multimodal data and the image data of any frame, the number of times any target is detected is determined, and the confidence feature information of any target in any frame is determined based on the determined number of times.

[0099] In this embodiment, a sliding time window can be used to obtain multiple frames of image data preceding any given frame, as well as the image data of any given frame, from multimodal data. The sliding step size is one frame. The duration of the sliding time window is usually shorter than the duration of a set time period, and can be set according to actual conditions.

[0100] Optionally, target detection is performed on multiple frames of image data preceding any frame in the multimodal data and on the image data of any frame, and the number of times any target is detected is determined. This includes: inputting multiple frames of image data preceding any frame in the multimodal data and on the image data of any frame into a target detection model for target detection, obtaining the target detection result output by the target detection model, and determining the number of times any target is detected based on the target detection result.

[0101] In this embodiment, the confidence feature information of target c in frame t is determined by the following formula. : ;Formula (5) in, Let c be the number of times target c is detected within the sliding time window. The higher the value, the more frequently target c appears. The higher; If the value is too low, the target c is considered to have disappeared. The clip function is used to restrict data to a specified range. The core function of the clip function is to restrict the element values ​​of an array, vector, or data frame to a specified interval. If the element value is less than the set lower limit, it is replaced with the lower limit value; if it is greater than the set upper limit, it is replaced with the upper limit value; values ​​within the interval remain unchanged.

[0102] In this embodiment, the velocity feature information of the dynamic feature information is determined in the following way: target detection is performed on image data of different frames to determine the center point coordinates of the target bounding box in different frames, and the velocity feature information of the dynamic feature information is determined based on the center point coordinates of the target bounding box in different frames and the second duration. The second duration is the duration between two adjacent frames, and the duration between adjacent frames (… The speed characteristic information depends on the set system frequency (frames per second, FPS). The speed characteristic information of dynamic feature information includes the speed characteristic information of each target in each frame.

[0103] In this embodiment, the velocity feature information of any target in any frame is determined as follows: target detection is performed on the image data of any frame and the image data of the previous frame of any frame; the coordinates of the first center point of the target bounding box of any target in any frame and the coordinates of the second center point of the target bounding box of any target in the previous frame of any frame are determined; and the velocity feature information of any target in any frame is determined based on the first center point coordinates, the second center point coordinates, and a second duration. The second duration is the duration between any frame and the previous frame of any frame.

[0104] Optionally, target detection is performed on image data of any frame and image data of the previous frame to determine the first center point coordinates of the target bounding box of any target in any frame and the second center point coordinates of the target bounding box of any target in the previous frame. This includes: inputting image data of any frame into a target detection model for target detection to obtain the target bounding box of any target in any frame and determining the first center point coordinates of the target bounding box of any target in any frame; inputting image data of the previous frame into a target detection model for target detection to obtain the target bounding box of any target in the previous frame and determining the second center point coordinates of the target bounding box of any target in the previous frame.

[0105] In this embodiment of the application, the velocity feature information of target c in frame t is determined by the following formula. : ;Formula (6) in, Let C be the coordinates of the center point of the target bounding box in frame t. For target c in The center point coordinates of the frame's target bounding box For t frames and The duration between frames, t frame and The frames are adjacent frames.

[0106] In this embodiment, the acceleration feature information of the dynamic feature information is determined as follows: based on the velocity feature information and the second duration, the acceleration feature information of the dynamic feature information is determined. The acceleration feature information of the dynamic feature information includes the acceleration feature information of each target in each frame.

[0107] In this embodiment, the acceleration feature information of any target in any frame is determined in the following way: based on the velocity feature information of any target in any frame, the velocity feature information of any target in the previous frame, and a second duration, the acceleration feature information of any target in any frame is determined. The second duration is the duration between any frame and the previous frame.

[0108] In this embodiment, the acceleration characteristic information of target c in frame t is determined by the following formula. : ;Formula (7) in, For the velocity feature information of target c in frame t, For target c in Frame velocity characteristics For t frames and The duration between frames.

[0109] In this embodiment, the rhythmic feature information of the dynamic feature information is determined by performing variance calculation on the velocity feature information. The rhythmic feature information of the dynamic feature information includes the rhythmic feature information of each target in each frame.

[0110] In this embodiment, the rhythmic feature information of any target in any frame is determined by the following method: variance calculation is performed on the velocity feature information of any target in each frame within a set time window to determine the rhythmic feature information of any target in any frame. The set time window is a sliding time window with a sliding step size of one frame. The duration of the set time window is shorter than the duration of a set time period. The latest time of the set time window is the time corresponding to any frame. The duration of the set time window can be set according to actual conditions.

[0111] In this embodiment, the rhythmic feature information of target c in frame t is determined by the following formula. : ;Formula (8) in, For target c in Frame velocity characteristics , To set the duration of the time window, It is a variance operation function. The goal is to characterize the rhythm of movement, i.e., rhythmic feature information, by calculating the variance of velocity.

[0112] In this embodiment of the application, in addition to baseline feature information, dynamic feature information and static feature information, frame feature information also includes noise suppression information. The noise suppression information is used to simulate and cancel small perturbations of detection / tracking noise, and can be modeled as zero-mean perturbations.

[0113] In this embodiment, the noise suppression information of the frame feature information is determined in the following way: based on the independent noise sampling information of multimodal data from different frames, the noise suppression information of the frame feature information is determined. The noise suppression information of the frame feature information includes the noise suppression information of each target in each frame, and the independent noise sampling information of each target in each frame follows a normal distribution.

[0114] This application can use the following formula to determine the noise suppression information of target c in frame t. : ;Formula (9) in, It follows a normal distribution, that is , It is the independent noise sampling information of target c in frame t. It is the independent noise sampling information of target c in frame t-1. It is the noise variance of the target c. In practice, it can also be simply implemented as a small-amplitude limited random quantity to break down the "rigidity" caused by perfect smoothing.

[0115] In this embodiment of the application, the target attention value of the target in the current frame is determined based on frame feature information. Figure 4 A flowchart illustrating a method for determining the target attention value of a target in the current frame, as provided in an embodiment of this application, is shown below. Figure 4 As shown, step 202 above includes at least the following steps 401-402: Step 401: Based on frame feature information, determine the attention value of the target in different frames.

[0116] In this embodiment of the application, determining the attention value of a target in different frames based on frame feature information includes: determining the initial attention value of the target in different frames based on frame feature information; and determining the attention value of the target in different frames based on the initial attention value of the target in different frames and a set smoothing coefficient. The frame feature information includes the frame feature information of each target in each frame.

[0117] Step 402: Determine the target attention value based on the target's attention values ​​in different frames and the total number of frames.

[0118] The total number of frames refers to the number of frames in the multimodal data, which corresponds to the duration of the set time period. The target attention value includes the target attention value of each target in the current frame.

[0119] In this embodiment of the application, the target attention value of any target in the current frame is determined based on frame feature information. Figure 5 A flowchart illustrating a method for determining the target attention value of any target in the current frame, as provided in an embodiment of this application, is shown below. Figure 5 As shown, it includes at least the following steps 501-503: Step 501: Based on the frame feature information of any target in any frame, determine the initial attention value of any target in any frame.

[0120] In this embodiment, determining the initial attention value of any target in any frame based on the frame feature information of any target in any frame includes: determining the static feature items of any target in any frame based on each static feature item and its corresponding weight; determining the dynamic feature items of any target in any frame based on each dynamic feature item and its corresponding weight; and determining the initial attention value of any target in any frame based on the static feature items, dynamic feature items, baseline feature information, and noise suppression information of any target in any frame. The weights corresponding to each static feature item and each dynamic feature item can be set according to actual conditions.

[0121] In this embodiment of the application, in order to suppress noise, the dynamic feature information of any target in any frame is processed by first-order exponential smoothing to obtain smoothed dynamic feature information.

[0122] In this embodiment, the smoothed dynamic feature information of target c in frame t is determined by the following formula. : ;Formula (10) in, It is a smoothing factor. It can be set according to the actual situation. The dynamic feature information of target c in frame t is the dynamic feature q. The dynamic feature information of target c after smoothing the dynamic features q in frame t-1.

[0123] In this embodiment, the static feature terms of target c in frame t are determined by the following formula. : ;Formula (11) Where j is a static feature information index, such as centrality, orientation, distance, confidence, etc. It is the weight of the static feature j of target c in scene s corresponding to frame t. The static feature information of target c in frame t is the static feature j.

[0124] In this embodiment, the dynamic feature terms of target c in frame t are determined by the following formula. : ;Formula (12) Where q is a dynamic feature index, such as velocity, acceleration, rhythm, etc. Let q be the weight of the dynamic feature q of target c in scene s corresponding to frame t. The dynamic feature information of target c in frame t after smoothing the dynamic features q.

[0125] In this embodiment, the initial attention value of target c in frame t is determined by the following formula. : ;Formula (13) in, This refers to the baseline feature information of target c in scene s corresponding to frame t. Let c be the static feature term of target c in frame t. Let c be the dynamic feature term of target c in frame t. The noise suppression information for target c in frame t.

[0126] Step 502: Based on the initial attention value of any target in any frame, the attention value of any target in the previous frame, and the set smoothing coefficient, determine the attention value of any target in any frame.

[0127] The smoothing coefficient is determined as follows: based on the scene corresponding to any frame and a second correspondence, the smoothing coefficient corresponding to the scene of that frame is determined. The second correspondence is the correspondence between each scene and each smoothing coefficient. This application can set the smoothing coefficient corresponding to each scene according to the actual situation. For example, the smoothing coefficient in the "interactive" scene is more "real-time", and the smoothing coefficient in the "office" scene is more stable.

[0128] In this embodiment, the attention value of target c in frame t is determined by the following formula. : ;Formula (14) in, Let be the smoothing coefficient corresponding to scene s, where scene s is the scene corresponding to frame t. Let c be the attention value for target c in frame t-1. Let be the initial attention value for target c in frame t.

[0129] Step 503: Based on the attention value of any target in each frame and the total number of frames, determine the target attention value of any target in the current frame.

[0130] In this embodiment of the application, determining the target attention value of any target in the current frame based on the attention value of any target in each frame and the total number of frames includes: determining a first sum of the attention values ​​of any target in each frame, and using the ratio of the first sum to the total number of frames as the target attention value of any target in the current frame.

[0131] In this embodiment of the application, the target attention value of target c in frame t can be determined by the following formula. : ;Formula (15) in, For target c in The attention value of the frame, set for a time period. , The range of values ​​is Let T be the duration of the set time period, and t be the current time.

[0132] In this embodiment of the application, the target attention value of any target in the current frame is determined based on the attention value of any target in each frame and the total number of frames. This includes: determining the target attention value of any target in the current frame based on the attention value of any target in each frame within a set backtracking time window and the total number of frames within the set backtracking time window.

[0133] Specifically, the last frame within the backtracking time window is set as the current frame; that is, the latest time of the backtracking time window is set as the current time. For example, the duration of the time window is... If the current frame is t, then the backtracking time window is set to... The duration of the backtracking time window is set to be shorter than the duration of the set time period. The duration of the backtracking time window can be set according to the actual situation; for example, the duration of the backtracking time window can be 2-3 seconds of physical time. The backtracking time window slides with the current frame t, with a step size of 1 frame, and is used to analyze recent attention patterns.

[0134] In this embodiment of the application, the target attention value of any target in the current frame is determined based on the attention value of any target in each frame within a set backtracking time window and the total number of frames within the set backtracking time window. This includes: determining a second sum of the attention values ​​of any target in each frame within the set backtracking time window, and using the ratio of the second sum to the total number of frames within the set backtracking time window as the target attention value of any target in the current frame.

[0135] In this embodiment of the application, the target attention value of target c in frame t can be determined by the following formula. : ;Formula (16) in, For target c in The attention value of the frame, with the backtracking time window set to... , The range of values ​​is Set the duration of the backtracking time window to be [duration]. The current time is t.

[0136] In this embodiment, based on the frame feature information of any target in any frame, an initial attention value for any target in any frame is determined. To suppress noise, the initial attention value of any target in any frame is smoothed using the attention value of the target in the previous frame and a set smoothing coefficient, thus obtaining the determined attention value for each target in any frame, thereby improving the accuracy of the attention value for each target in each frame. Furthermore, this application determines the target attention value for any target in the current frame based on the attention values ​​of any target in each frame and the total number of frames within a set time period, thereby improving the accuracy of the target attention value.

[0137] In this embodiment of the application, step 203 above determines the stimulus event of the current frame based on multimodal data, including some or all of the following methods: The first method determines the burst level information of the current frame based on frame feature information. If the burst level information of the current frame exceeds the set level threshold, a stimulus event is generated.

[0138] The level threshold can be set according to the actual situation. For example, if the suddenness level information is R0, it means that the event cannot change the emotional state or the change in the emotional state is small; therefore, the level threshold is set to R0.

[0139] Specifically, based on frame feature information, the burst level information of the current frame is determined, including: based on the frame feature information of each target in the current frame and the frame feature information of each target in the previous frame, the burst level information of the current frame is determined.

[0140] The frame feature information of each target in the current frame includes static feature information and dynamic feature information, namely, some or all of the centrality feature information, orientation feature information, distance feature information, confidence feature information, velocity feature information, acceleration feature information and rhythm feature information.

[0141] In this embodiment of the application, the burst level information of the current frame is determined based on the frame feature information of each target in the current frame and the frame feature information of each target in the previous frame. Figure 6 A flowchart illustrating a method for determining burst level information of the current frame, as provided in an embodiment of this application, is shown below. Figure 6 As shown, it includes at least the following steps 601-602: Step 601: For any target, based on the frame feature information of any target in the current frame, the frame feature information of any target in the previous frame, and the second duration, determine the feature change rate of any target in the current frame.

[0142] The second duration is the time interval between the current frame and the previous frame. The feature change rate of any target in the current frame includes the feature change rates of feature information from each frame.

[0143] In this embodiment of the application, determining the feature change rate of any target in the current frame based on the frame feature information of any target in the current frame, the frame feature information of any target in the previous frame of the current frame, and the second duration includes: for any frame feature information, determining the feature change rate of any frame feature information of any target in the current frame based on any frame feature information of any target in the current frame, any frame feature information of any target in the previous frame of the current frame, and the second duration.

[0144] In this embodiment, the feature change rate of target c at feature information index m in frame t is determined by the following formula. : ;Formula (17) in, The frame feature information of target c in frame t is indexed by the frame feature information of m. For target c in The frame feature information index m of the frame. For t frames and The duration between frames. m is an index of frame feature information, such as centrality, orientation, distance, confidence, velocity, acceleration, and rhythm.

[0145] Step 602: Based on the characteristic change rate of each target in the current frame, determine the burst index value of the current frame, and determine the burst level information corresponding to the set index value range to which the burst index value of the current frame belongs, so as to obtain the burst level information of the current frame.

[0146] In this embodiment, determining the burst index value of the current frame based on the feature change rate of each target in the current frame includes: determining the burst index value of the current frame based on the feature change rate of each target's feature information in each frame and its corresponding weight. The weights corresponding to the feature change rates of each target's feature information in each frame can be set according to actual conditions.

[0147] In this embodiment, the burst index value of frame t can be determined using the following formula. : ;Formula (18) in, Let m be the feature change rate of target c in frame t, which is the index of frame feature information. This represents the absolute value of the characteristic rate of change. for The corresponding weights are the sensitivity weights for sudden responses. For example, a sudden decrease in hand distance has a high weight, while a slight change in facial confidence has a low weight. C is the target set, and M is the feature index set, including some or all of centrality, orientation, distance, confidence, velocity, acceleration, and rhythm.

[0148] In this embodiment of the application, multiple preset index value ranges are set in the following manner: a set of index thresholds are set. Based on this set of indicator thresholds, six preset indicator value ranges are defined, and the corresponding outbreak level information is set for each preset indicator value range. The first preset indicator value range is... The corresponding emergency level information is R0; the range of the second set indicator value is... The corresponding emergency level information is R1; the range of the third set indicator value is... The corresponding emergency level information is R2; the range of the fourth set indicator value is... The corresponding emergency level information is R3; the range of the 5th set indicator value is... The corresponding emergency level is R4; the range of the 6th set indicator value is... The corresponding emergency level information is R5.

[0149] In this embodiment, the burst index value of frame t is determined by the following formula. The information on the emergency level corresponding to the set indicator value range. : ;Formula (19) in, , , , and For the set indicator threshold, It can be set according to the actual situation. R0, R1, R2, R3, R4 and R5 are the emergency level information.

[0150] In this embodiment of the application, during the process of determining the burst index value of the current frame, other frame feature information can be superimposed to further correct the burst level information, such as the hand continuously approaching from the baseline distance to the ultra-close distance, and the speed and acceleration increasing simultaneously.

[0151] In this embodiment of the application, if the burst level information of the current frame is R4 or R5, then in the continuous (Cooldown time) Ignore R0 to R3 within the frame, and only retain the information of the highest level of burst.

[0152] The second approach determines the duration of the target's primary attention based on the target's attention values ​​in different frames. If the duration of the target's primary attention exceeds a first duration threshold, a stimulus event is generated.

[0153] Specifically, the target corresponding to the maximum attention value of each target in the same frame is taken as the main attention target of the frame, and the duration of each target as the main attention target is determined based on the main attention target of each frame. If the duration of at least one target as the main attention target exceeds the first duration threshold, a stimulus event is generated.

[0154] The first duration threshold can be set according to the actual situation. The first duration threshold is the maximum duration during which the change in emotional state has a relatively small impact when the target is the primary focus of attention.

[0155] In this embodiment of the application, the main attention target of frame t is determined by the following formula. : ;Formula (20) in, Let C be the attention value of target c in frame t, where C is the set of targets, and argmax() is the function that takes the maximum value. Main attention target. When a switch occurs, the "state retention counter" of some features is reset to avoid false bursts during the switch. For example, if the face is continuously focused on as the primary target for 100 frames, the state retention counter is recorded as 100. When the primary target for attention switches to the hand, the state retention counter needs to be reset.

[0156] The third approach involves performing emotion recognition on the speech data of multimodal data from different frames, determining the duration of speech data corresponding to different emotions based on the emotion recognition results, and generating a stimulus event if an emotion exceeds a second duration threshold.

[0157] The second duration threshold can be set according to the actual situation. The second duration threshold is the maximum duration for which the voice data corresponding to each emotion has a small impact on the change of emotional state.

[0158] In this embodiment, emotion recognition is performed on speech data from multimodal data of different frames, and the duration of speech data corresponding to different emotions is determined based on the emotion recognition results. This includes: inputting speech data from different frames into an emotion recognition model for emotion recognition, obtaining the emotion recognition results output by the model, and determining the duration of speech data corresponding to different emotions based on the emotion recognition results. The emotion recognition model can be pre-trained using machine learning methods (such as supervised learning methods). The base model used to train the emotion recognition model can be various models with recognition functions, such as convolutional neural network models, neural network models, etc.

[0159] The fourth method involves performing target detection on image data from multimodal data in different frames, determining the first duration corresponding to image data that does not include the target based on the target detection results, and generating a stimulus event if the first duration exceeds the third duration threshold.

[0160] The third duration threshold can be set according to the actual situation. The third duration threshold is the maximum duration during which image data excluding the target (the target leaving the detection range) has little impact on changes in emotional state.

[0161] In this embodiment of the application, target detection is performed on image data of multimodal data of different frames, and a first duration corresponding to image data that does not include the target is determined based on the target detection result. This includes: inputting image data of different frames into a target detection model for target detection, obtaining the target detection result output by the target detection model, and determining the first duration corresponding to image data that does not include the target based on the target detection result.

[0162] The fifth method is to determine the duration of touch on different parts of the robot based on pressure data from multimodal data in different frames. If the duration exceeds the fourth duration threshold, a stimulus event is generated.

[0163] The pressure data in the multimodal data of different frames consists of pressure data collected by pressure sensors on various parts of the robot. The fourth duration threshold can be set according to the actual situation, while the third duration threshold is the maximum duration for which the user's touch on the robot's parts has a relatively small impact on changes in emotional state.

[0164] In this embodiment, the stimulus events of the current frame are determined based on multimodal data, and an external stimulus set is constructed. Let i be the total number of stimulus events. Stimulus set The stimulus events can include burst level information from t-frames in the attention layer. The duration of attention focused on the face The duration of attention focused on the Hand The duration of voice data corresponding to user laughter from speech / sound. The duration of voice data corresponding to the user's oppressive or angry tone. The duration of image data from other sensors that does not include the user (when the user leaves the detection range). The duration of user touching the robot's head .

[0165] In this embodiment of the application, when the current time is within the execution time period of the task, the first emotional state information is determined based on the duration corresponding to the execution time period, the third duration, the initial emotional state information corresponding to the execution time period, and the termination emotional state information. The emotional state information is then determined based on the first emotional state information and the second emotional state information determined using the weight of the stimulus event. Figure 7 A flowchart illustrating a method for determining the emotional state information of the current frame, as provided in an embodiment of this application, is shown below. Figure 7 As shown, step 204 above includes at least the following steps 701-704: Step 701: If the current time is within the task's execution time period, then determine the initial emotional state information and the termination emotional state information for the execution time period.

[0166] In this embodiment, each task corresponds to an initial emotional state and a termination emotional state. The initial and termination emotional state information for each task can be set according to actual conditions. Each task corresponds to an execution duration. If a task starts execution, the execution time period is determined based on the task's start time and execution duration.

[0167] Step 702: Determine the first emotional state information based on the duration corresponding to the execution time period, the third duration, the initial emotional state information, and the termination emotional state information.

[0168] The third duration is determined based on the earliest time of the current time and the execution time period. The initial emotional state information is the emotional state information corresponding to the earliest time of the task's execution time period, and the termination emotional state information is the emotional state information corresponding to the latest time of the task's execution time period. Specifically, when the robot performs a task, it is desirable for the emotion to smoothly transition from the initial emotional state information to the termination emotional state information within the task's execution time period. For example, if the current task performed by the robot is "to greet others warmly," it is desirable for the emotion to smoothly transition from the initial emotional state information to the termination emotional state information within the task's execution time period. linear transition to The execution time for this task is 2 seconds. If the current task the robot is performing is "quiet companionship," then we hope to... Slow transition to The execution time of this task is 8 seconds.

[0169] In this embodiment, the first emotional state information of frame t is determined by the following formula. : ;Formula (21) in, This is information about the initial emotional state. To terminate emotional state information, This is the earliest time during which the task will be executed. This is the latest time during the task's execution period. The current time is [time], and the task execution time period is [duration]. , .

[0170] Step 703: Determine the second emotional state information based on the weight of the stimulus event.

[0171] In this embodiment of the application, a corresponding weight is set for each stimulus event, which can be set according to the actual situation.

[0172] Specifically, for the i-th stimulus event, its weight is defined. : ;Formula (22) in, Let V be the weight of the i-th stimulus event under the pleasure level V, that is, the influence of the i-th stimulus event on the pleasure level V. Let be the weight of the i-th stimulus event under arousal level A, that is, the influence of the i-th stimulus event on arousal level A. Let be the weight of the i-th stimulus event under the degree of domination D, that is, the influence of the i-th stimulus event on the degree of domination D.

[0173] For example, in response to stimulating events (Moderately sudden) This stimulus event The corresponding weights can be set to This primarily aims to increase arousal levels. It targets stimulating events. That is, the stimulus event is when the user actively approaches and smiles. The corresponding weights can be set to This increases pleasure and slightly arousal. (Targeting stimulating events.) That is, the user uses a strong command tone when addressing the stimulus event. The corresponding weights can be set to It reduces pleasure and a slight degree of dominance, while increasing arousal.

[0174] In this embodiment, determining the second emotional state information based on the weights of the stimulus events includes: determining the second emotional state information based on the sum of the weights of each stimulus event. The second emotional state information can reflect the emotional state increment caused by external stimuli in frame t.

[0175] In this embodiment, the second emotional state information of frame t is determined by the following formula. : ;Formula (23) in, Let be the weight of the i-th stimulus event. For a set of stimulating events, Let i be the i-th stimulus event in the set of stimulus events.

[0176] Step 704: Determine the emotional state information based on the first emotional state information and the second emotional state information.

[0177] In this embodiment, the emotional state information of frame t is determined by the following formula. : ;Formula (24) in, The first emotional state information of frame t. For the second emotional state information of frame t, clip() is a function that restricts the value of each element in the vector to the range [-1, 1].

[0178] In this embodiment of the application, if the second emotional state information determined by the weight of the stimulus event is less than a set threshold when the current time is not within the execution time period of the task, then the fourth emotional state information is determined based on the third emotional state information when the robot is first detected to be in an idle state, the set emotional state information, and the fourth duration, and the emotional state information is determined based on the second emotional state information and the fourth emotional state information. Figure 8 A flowchart illustrating another method for determining the emotional state information of the current frame provided in this application embodiment is shown below. Figure 8 As shown, step 204 above includes at least the following steps 801-804: Step 801: If the current time is not within the task's execution time period, determine the second emotional state information based on the weight of the stimulus event.

[0179] The specific process of determining the second emotional state information based on the weight of the stimulus event in step 801 is the same as the specific process in step 703, and will not be repeated here.

[0180] Step 802: Compare the second emotional state information with the set threshold to obtain the second comparison result.

[0181] The threshold can be set according to the actual situation.

[0182] Step 803: If the second comparison result indicates that the second emotional state information is less than the set threshold, then the robot is determined to be in an idle state, and the fourth emotional state information is determined based on the third emotional state information when the robot is first detected to be in an idle state, the set emotional state information, and the fourth duration.

[0183] The fourth duration is determined based on the current time and the time when the robot was first detected to be in an idle state. Emotional state information can be set according to actual circumstances; for example, setting emotional state information... .

[0184] In this embodiment, the fourth emotional state information is determined based on the third emotional state information when the robot is first detected to be in an idle state, the set emotional state information, and the fourth duration. This includes: determining the fourth emotional state information based on the third emotional state information, the set emotional state information, the set duration corresponding to the dissipation period, and the fourth duration. The dissipation period is the longest time required for an emotional state to completely "cool down" back to a neutral baseline from any "excited" state without any new stimuli.

[0185] In this embodiment, the fourth emotional state information of frame t is determined by the following formula. : ;Formula (25) in, To set emotional state information, for The third emotional state information of the frame, The time when the robot was first detected to be in an idle state is the time of the last non-dissipative update. This is the fourth duration. To set the duration corresponding to the dissipation period, the above formula guarantees... Time remains unchanged. Completely return to the set emotional state information .

[0186] Step 804: Determine the emotional state information based on the second and fourth emotional state information.

[0187] In this embodiment, the emotional state information of frame t is determined by the following formula. : ;Formula (26) in, This refers to the fourth emotional state information in frame t. For the second emotional state information of frame t, clip() is a function that restricts the value of each element in the vector to the range [-1, 1].

[0188] In this embodiment of the application, after comparing the second emotional state information with a set threshold in step 802 to obtain a second comparison result, the method further includes: if the second comparison result indicates that the second emotional state information is greater than or equal to the set threshold, then the second emotional state information is used as the emotional state information.

[0189] In this embodiment, the stimulus event of the current frame is determined based on multimodal data, and the emotional state information of the current frame is determined based on the weight of the stimulus event. Figure 9 A detailed flowchart illustrating a method for determining the emotional state information of the current frame, as provided in this application embodiment, is shown below. Figure 9 As shown, it includes at least the following steps 901-911: Step 901: Determine the stimulus event of the current frame based on the multimodal data.

[0190] Step 902: Determine if the current time is within the task's execution time period. If yes, proceed to step 903; otherwise, proceed to step 907.

[0191] Step 903: Determine the initial emotional state information and the termination emotional state information for the execution time period.

[0192] Step 904: Determine the first emotional state information based on the duration corresponding to the execution time period, the third duration, the initial emotional state information, and the termination emotional state information.

[0193] Step 905: Determine the second emotional state information based on the weight of the stimulus event.

[0194] Step 906: Determine the emotional state information based on the first emotional state information and the second emotional state information.

[0195] Step 907: Determine the second emotional state information based on the weight of the stimulus event.

[0196] Step 908: Determine whether the second emotional state information is less than the set threshold. If yes, proceed to step 909; otherwise, proceed to step 911.

[0197] Step 909: Determine that the robot is in an idle state, and determine the fourth emotional state information based on the third emotional state information when the robot is first detected to be in an idle state, the set emotional state information, and the fourth duration.

[0198] Step 910: Determine the emotional state information based on the second and fourth emotional state information.

[0199] Step 911: Use the second emotional state information as the emotional state information.

[0200] In this embodiment, step 205, based on the scene tendency value, target attention value, and emotional state information of the previous frame, determines the target scene tendency value of the current frame. This includes: determining the target scene tendency value based on the scene tendency value, target attention value, emotional state information, and their respective weights from the previous frame. Each weight can be set according to actual conditions, and the target attention value includes the target attention value of each target in the current frame.

[0201] In this embodiment of the application, the target scene tendency value of frame t can be determined using the following formula. : ;Formula (27) in, for Scene bias value of the frame, Let c be the target attention value in frame t. The emotional state information for frame t. for The duration between frame 1 and frame t. Frame t is adjacent to frame t, and C is the target set. The forgetting coefficient, i.e. The corresponding weights are used to control the degree to which historical evidence is preserved. The value range is [0,1]. for The corresponding weights for The corresponding weights. , and It can be set according to the actual situation.

[0202] In this embodiment of the application, if the target set C includes hands and faces, the target scene tendency value of frame t can be determined using the following formula. : ;Formula (28) in, The target attention value of the hand in frame t. Let be the target attention value for the face in frame t. for Scene bias value of the frame, The emotional state information for frame t. for The duration between frame 1 and frame t. This represents the forgetting coefficient. for The corresponding weight, namely the support weight of "hand being watched" for interactive scenarios. for The corresponding weight, namely the support weight of "face being steadily focused on" for the office scenario (represented by a negative sign in the formula). for The corresponding weights represent the gain of Arousal (motivation) in the emotion on the willingness to interact in the scenario. , , and It can be set according to the actual situation.

[0203] In this embodiment of the application, before step 205, which determines the target scene tendency value of the current frame based on the scene tendency value, target attention value, and emotional state information of the previous frame, the method further includes: determining target outbreak level information based on frame feature information. The target outbreak level information includes outbreak level information from different frames.

[0204] In this embodiment of the application, the target burst level information is determined based on frame feature information. Figure 10 A flowchart illustrating a method for determining target outbreak level information provided in this application embodiment is shown below. Figure 10 As shown, it includes at least the following steps 1001-1002: Step 1001: Based on frame feature information and second duration, determine the feature change rate of the target in different frames.

[0205] Step 1002: Determine the target burst level information based on the target's feature change rate in different frames.

[0206] Specifically, based on the characteristic change rate of the target in different frames, the burst level information of the target is determined, including: determining the burst index value of different frames based on the characteristic change rate of the target in different frames; determining the burst level information of different frames based on the burst level information corresponding to the set index value range to which the burst index value of different frames belongs; determining the number of frames in which the burst level information exceeds the set burst level threshold, and thus obtaining the target burst level information.

[0207] In this embodiment of the application, determining the number of frames in which the burst level information exceeds a set burst level threshold to obtain the target burst level information includes: determining the number of frames in which the burst level information exceeds a set burst level threshold within a time window to obtain the target burst level information. The latest time of the time window is the current time, and the duration of the time window is less than or equal to the duration of a set time period. The duration of the time window can be set according to the actual situation.

[0208] In this embodiment of the application, the target burst level information of frame t can be determined using the following formula. : ;Formula (29) in, for Frame burst level information, R3 sets the burst level threshold, and the time window is... , This represents the duration of the time window.

[0209] In this embodiment of the application, the burst level information of any frame is determined by the following method: based on the frame feature information of each target in any frame and the frame feature information of each target in the previous frame of any frame, the burst level information of any frame is determined. Figure 11 A flowchart illustrating a method for determining burst level information of any frame, as provided in this application embodiment, is shown below. Figure 11 As shown, it includes at least the following steps 1101-1102: Step 1101: For any target, based on the frame feature information of any target in any frame, the frame feature information of any target in the previous frame and the second duration, determine the feature change rate of any target in any frame.

[0210] The second duration is the duration between any given frame and the frame preceding it. The specific process of step 1101 is the same as that of step 601 above, and will not be described in detail here.

[0211] Step 1102: Based on the characteristic change rate of each target in any frame, determine the burst index value of any frame, and determine the burst level information corresponding to the set index value range to which the burst index value of any frame belongs, so as to obtain the burst level information of any frame.

[0212] The specific process of step 1102 is the same as that of step 602 above, and will not be described in detail here.

[0213] In this embodiment, based on the frame feature information of any target in any frame, the frame feature information of any target in the previous frame, and the second duration, the feature change rate of any target in any frame is determined to reflect the changes in frame feature information within the second duration. Furthermore, based on the feature change rate of each target in any frame, this application determines the burst index value of any frame and the burst level information corresponding to the set index value range to which the burst index value of any frame belongs, thereby improving the accuracy of the burst level information.

[0214] In this embodiment of the application, if the target outbreak level information is determined before determining the target scene tendency value of the current frame, the target scene tendency value of the current frame is determined by the following method: the target scene tendency value is determined based on the scene tendency value of the previous frame, the target attention value, the target outbreak level information and the emotional state information of the current frame.

[0215] In this embodiment, the target scene tendency value of the current frame is determined based on the scene tendency value of the previous frame, the target attention value of each target in the current frame, the burst level information of each frame, and the emotional state information of the current frame. This includes determining the target scene tendency value based on the scene tendency value of the previous frame, the target attention value, the target burst level information, the emotional state information, and their respective weights. Each weight can be set according to actual conditions.

[0216] In this embodiment of the application, the target scene tendency value of frame t can be determined using the following formula. : ;Formula (30) in, for Scene bias value of the frame, Let c be the target attention value in frame t. The emotional state information for frame t. For the target burst level information of frame t, for The duration between frame 1 and frame t. Frame t is adjacent to frame t, and C is the target set. This represents the forgetting coefficient. for The corresponding weights for The corresponding weights for The corresponding weights. , , and It can be set according to the actual situation.

[0217] In this embodiment of the application, if the target set C includes hands and faces, the target scene tendency value of frame t can be determined using the following formula. : ;Formula (31) in, The target attention value of the hand in frame t. Let be the target attention value for the face in frame t. for Scene bias value of the frame, The emotional state information for frame t. For the target burst level information of frame t, for The duration between frame 1 and frame t. Frame t is adjacent to frame t, and C is the target set. This represents the forgetting coefficient. for The corresponding weights for The corresponding weights, i.e., the bias of the target burst level information in frame t towards the interactive scene, for The corresponding weights. , , , and It can be set according to the actual situation.

[0218] In this embodiment, based on the multimodal data of the target in different frames, the target attention value, target outbreak level information and emotional state information of the target in the current frame are determined, and based on the scene tendency value, target attention value, target outbreak level information and emotional state information of the previous frame of the current frame, the target scene tendency value is determined. Figure 12 A detailed flowchart illustrating another scene switching method for a robot provided in this application embodiment is shown below. Figure 12 As shown, it includes at least the following steps 1201-1210: Step 1201: Extract features from the multimodal data of the target in different frames to obtain frame feature information.

[0219] Step 1202: Determine the target attention value of the target in the current frame based on frame feature information.

[0220] Step 1203: Determine the stimulus event of the current frame based on multimodal data.

[0221] Step 1204: Determine the emotional state information of the current frame based on the weight of the stimulus event.

[0222] Step 1205: Determine the target burst level information based on frame feature information.

[0223] In this embodiment, the cerebellar attention module in the robot is used to determine the target burst level information based on frame feature information.

[0224] Step 1206: Determine the target scene tendency value of the current frame based on the scene tendency value, target attention value, target burst level information and emotional state information of the previous frame of the current frame.

[0225] Step 1207: Select the scene corresponding to the set tendency value range to which the target scene tendency value belongs as the target scene.

[0226] Step 1208: Determine whether the target scene and the current scene are the same. If they are the same, proceed to step 1209; otherwise, proceed to step 1210.

[0227] Step 1209: Keep the current scene unchanged.

[0228] Step 1210: Switch the current scene to the target scene.

[0229] In this embodiment, in step 206 or step 1207 above, the scene corresponding to the set tendency value range to which the target scene tendency value belongs is taken as the target scene. Each set tendency value range corresponds to one scene. This application divides the scene into set tendency value ranges based on multiple set tendency thresholds. Each tendency threshold can be set according to actual conditions. The brain decision-making module in the robot is used to execute steps 1206-1210 above.

[0230] For example, the robot has two scenarios: an interactive scenario and an office scenario. An upper threshold is set for entering the interactive scenario. And the lower threshold for entering the office environment. , The target scene of frame t is determined by the following formula. : ;Formula (32) in, The target scene tendency value for frame t. The target scene is frame t-1. Specifically, when the combined evidence of "hand + burst + arousal" (target scene tendency value) is consistently high, the target scene is determined to be an interactive scene, that is, "high attention to hand / object + high frequency burst + relatively high arousal" is regarded as evidence of a bias towards an interactive scene; when the combined evidence of "facial steady state + low burst" (target scene tendency value) is high, the target scene is determined to be an office scene, that is, "facial steady state attention + low burst + stable emotion" is regarded as evidence of a bias towards an office scene. This application maps these signals into a continuous "target scene tendency value" through the above formula (30) and switches scenes through threshold hysteresis rules.

[0231] In this embodiment of the application, after determining whether to switch scenes based on the first comparison result, the method further includes: determining the main attention target of the current frame based on the attention value of the target in the current frame; and selecting and executing a target task chain from multiple preset task chains based on the target scene, emotional state information, the main attention target of the current frame, and historical task information.

[0232] Historical task information includes task execution records within a historical time period, serving as the robot's short-term working memory. Its core function is to avoid repetitive behaviors and prevent action conflicts. It records recently executed task chains and sets cooldown times for similar behaviors; simultaneously, it tracks the status of currently executing actions to ensure smooth transitions to new task chains, making interactions more natural and logical. The latest time within a historical time period is the current time, and the duration of each historical time period can be set according to actual needs.

[0233] Specifically, determining the main attention target of the current frame based on the attention value of the target in the current frame includes: taking the target corresponding to the maximum value among the attention values ​​of each target in the current frame as the main attention target of the current frame. In the embodiments of this application, the main attention target of the current frame can be determined using the above formula (20).

[0234] In this embodiment, the task chain can be viewed as an ordered set of "atomic behavioral states". The i-th task includes the behavior type (e.g., gazing, greeting, imitating actions, standby, etc.); the expected duration range; the start and end conditions; and optional interruption conditions (e.g., high-level emergency response).

[0235] In this embodiment, the robot's brain decision-making module updates the task each time it needs to, based on the target scenario. Emotional state information of the current frame The main attention target of the current frame It also includes historical task information, allowing you to select and execute the target task chain from multiple preset task chains.

[0236] For example, in an "office" scenario where attention is focused on a face, a low-motion task chain of "slight eye contact + nodding response" might be chosen; in an "interactive" scenario where hands and objects repeatedly receive high attention, a game-like task chain of "following hand movement + imitating actions + verbal feedback" would be chosen. Specifically, if the target scenario is "interactive" and the primary focus of attention is on a hand / object, a task chain of "following hand / toy" should be prioritized; if the target scenario is "office" and the primary focus of attention is on a face, and valence is high, a "companionship" task chain should be chosen; if R4 / R5 occur consecutively and a significant increase in arousal is observed, a "soothing / security check" type task chain can be initiated from any scenario.

[0237] When the target task chain Once selected, from the task Start executing sequentially, each task Internal execution: The expected execution duration of the task is mapped to the execution time period of the task in step 701 above, which is used to update the emotional state information; and, for each frame within the execution time period, the following check is performed: whether it meets the requirements. The completion conditions (such as time limit achievement or positive user feedback); whether a high-level emergency (such as R4 or R5) has occurred, requiring the termination of the target task chain and switching to a security-related task chain. Enter after completion This continues until the target task chain ends or is interrupted.

[0238] In this embodiment, the robot's brain decision-making module can reassess the scene and emotional state at each stage of task execution and switch to a new task chain when necessary. If the main attention target changes significantly during the execution of the task chain (e.g., switching from the hand back to the face), or the emotion changes from positive to negative, the brain decision-making module can end the current behavior early and switch to a more suitable task chain (e.g., switching from a "play" type task chain to an "observation + voice reassurance" type task chain).

[0239] In this embodiment, the robot's decision-making module issues the following set of template parameters for each scenario s. : ;Formula (33) in, This refers to the baseline feature information of target c in scene s corresponding to frame t. It is the weight of the static feature j of target c in scene s corresponding to frame t. Let q be the weight of the dynamic feature q of target c in scene s corresponding to frame t. is the smoothing coefficient corresponding to scene s.

[0240] These parameters are fed into the cerebellar attention module for actual attention value calculation and target switching logic, enabling the same algorithm to exhibit different styles and sensitivities in different scenarios.

[0241] In this embodiment, the robot includes a cerebellar attention module, an emotion model module, and a brain decision-making module. The cerebellar attention module performs target-level attention scoring on multimodal data at frame-level frequency, determines the target attention value of the target in the current frame, detects short-term mutations in the attention value, and generates hierarchical burst level information. The emotion model module is used to detect the stimulus events activated in the current frame and obtains emotional state information using VAD three-dimensional representation based on the weights of the stimulus events. The brain decision-making module reads the following signal: from the cerebellar attention module: the target attention value of target c in frame t. The primary attention target g(t) of frame t; the burst level information R(t) and target burst level information of frame t. From the emotion model module: Emotional state information of frame t. Based on the scene tendency value of the previous frame, the target attention value of target c in frame t, the main attention target in frame t, the burst level information of frame t, the target burst level information, historical task information, and the emotional state information of frame t, scene judgment and task chain decision are made, and scene template parameters are sent to the cerebellar attention module. In this application, the three layers jointly generate a unified control signal to drive the robot's multimodal output, including facial expressions, actions, and speech, to achieve biomimetic and natural human-computer interaction. Moreover, the three layers share a potential control vector, making speech, facial expressions, and postures coordinated and consistent. The cerebellar attention module, emotion model module, and brain decision-making module in the robot are independent of each other, which facilitates expansion and deployment in different robot forms and interaction scenarios.

[0242] In this application embodiment, feature extraction is performed on the multimodal data of the target in different frames to obtain frame feature information. Frame feature information includes baseline feature information, static feature information, and dynamic feature information. Based on the frame feature information, the target attention value in the current frame is determined, improving the accuracy and robustness of the target attention value calculation. This application determines the stimulus events of the current frame based on multimodal data and, based on the weights of the stimulus events, determines the emotional state information of the current frame represented by the VAD three-dimensional state. This allows for smooth and interpretable emotion updates through the VAD three-dimensional state and the stimulus event set, improving the accuracy of the emotional state information of the current frame. Based on the frame feature information, this application determines the target's abruptness level information, thereby achieving a millisecond-level response to abrupt stimuli and maintaining stability through a smoothing mechanism. This application determines the target scene tendency value of the current frame based on the scene tendency value, target attention value, target burst level information, and emotional state information of the previous frame of the current frame, thereby reflecting the scene bias of the current frame. Furthermore, it determines the target scene based on the scene corresponding to the set tendency value range to which the target scene tendency value belongs. Based on the target scene and the current scene, it determines whether to switch scenes, thereby improving the coherence of scene switching of the robot, making scene switching more natural and having "intentional" characteristics, and improving the user experience.

[0243] Based on the same inventive concept, this application provides a schematic diagram of the structure of a scene switching device for a robot, as shown below. Figure 13 As shown, the scene switching device of the robot includes: The first determining module 131 is used to extract features from the multimodal data of the target in different frames to obtain frame feature information, and to determine the target attention value of the target in the current frame based on the frame feature information. The second determining module 132 is used to determine the stimulus event of the current frame based on the multimodal data, and to determine the emotional state information of the current frame based on the weight of the stimulus event, wherein the stimulus event is an event that causes a change in the emotional state. The third determining module 133 is used to determine the target scene tendency value of the current frame based on the scene tendency value of the previous frame of the current frame, the target attention value and the emotional state information. The switching module 134 is used to determine the target scene based on the target scene tendency value, compare the target scene with the current scene, and determine whether to switch scenes based on the first comparison result.

[0244] Optionally, the first determining module 131 is used to: Based on the frame feature information, the attention value of the target in different frames is determined; The target attention value is determined based on the target's attention values ​​in different frames and the total number of frames.

[0245] Optionally, the second determining module 132 is used to perform some or all of the following steps: Based on the frame feature information, the burst level information of the current frame is determined. If the burst level information of the current frame exceeds the set level threshold, the stimulus event is generated. Based on the attention values ​​of the target in different frames, the duration of the target as the primary attention target is determined. If the duration of the target as the primary attention target exceeds a first duration threshold, the stimulus event is generated. Emotion recognition is performed on the speech data of multimodal data from different frames. Based on the emotion recognition results, the duration of the speech data corresponding to different emotions is determined. If there is an emotion that exceeds a second duration threshold, the stimulus event is generated. Target detection is performed on image data of multimodal data from different frames. Based on the target detection results, a first duration corresponding to image data that does not include the target is determined. If the first duration exceeds a third duration threshold, the stimulus event is generated. Based on the pressure data from multimodal data of different frames, the duration of touch on different parts of the robot is determined. If the duration exceeds the fourth duration threshold, the stimulus event is generated.

[0246] Optionally, before determining the target scene tendency value of the current frame, the third determining module 133 is further configured to: Based on the frame feature information, the target burst level information is determined; The third determining module 133 is used for: The target scene tendency value is determined based on the scene tendency value of the previous frame of the current frame, the target attention value, the target suddenness level information, and the emotional state information.

[0247] Optionally, the third determining module 133 is used to: Based on the frame feature information and the second duration, the feature change rate of the target in different frames is determined, where the second duration is the duration between two adjacent frames; The burst level information of the target is determined based on the characteristic change rate of the target in different frames.

[0248] Optionally, the second determining module 132 is used to: If the current time is within the task's execution time period, then determine the initial emotional state information and the termination emotional state information for the execution time period. Based on the duration corresponding to the execution time period, the third duration, the initial emotional state information, and the termination emotional state information, the first emotional state information is determined, and the third duration is determined based on the current time and the earliest time of the execution time period; Based on the weights of the stimulus events, the second emotional state information is determined; The emotional state information is determined based on the first emotional state information and the second emotional state information.

[0249] Optionally, the second determining module 132 is used to determine the first emotional state information using the following formula. : ; in, This is information about the initial emotional state. To terminate emotional state information, This refers to the earliest time of the task's execution period. This refers to the latest time during which the task will be executed. For the current time, .

[0250] Optionally, the second determining module 132 is used to: If the current time is not within the task's execution time period, the second emotional state information is determined based on the weight of the stimulus event. The second emotional state information is compared with a set threshold to obtain a second comparison result; If the second comparison result indicates that the second emotional state information is less than a set threshold, then it is determined that the robot is in an idle state, and based on the third emotional state information when the robot is first detected to be in an idle state, the set emotional state information, and the fourth duration, the fourth emotional state information is determined, wherein the fourth duration is determined based on the current time and the time when the robot is first detected to be in an idle state; The emotional state information is determined based on the second emotional state information and the fourth emotional state information.

[0251] Optionally, after comparing the second emotional state information with a set threshold to obtain a second comparison result, the second determining module 132 is further configured to: If the second comparison result indicates that the second emotional state information is greater than or equal to a set threshold, then the second emotional state information is used as the emotional state information of the current frame.

[0252] Optionally, the first determining module 131 is used to determine the frame feature information by the following method: Based on the scene corresponding to different frames, the target and the first correspondence relationship, the baseline feature information of the frame feature information is determined, and the first correspondence relationship is the correspondence relationship between each scene, each target and each baseline feature information. Static feature extraction is performed on the image data of multimodal data from different frames to determine the static feature information of the frame feature information; Dynamic feature extraction is performed on the image data of multimodal data from different frames to determine the dynamic feature information of the frame feature information.

[0253] Optionally, the first determining module 131 is used to perform some or all of the following steps: Target detection is performed on image data from different frames to determine the target bounding boxes of the target in different frames. Based on the center point coordinates of the target bounding boxes in different frames and the center point coordinates of the image data in different frames, the centrality feature information of the static feature information is determined. Based on image data from different frames, a deep learning method is used to determine the target angle of the target in different frames, and based on the target angle of the target in different frames, the orientation feature information of the static feature information is determined, wherein the target angle is the angle between the target normal and the camera's line of sight; Target detection and key point detection are performed on image data of different frames to determine the key points of the target in different frames. Based on the key points of the target in different frames and the depth image data of multimodal data in different frames, the distance feature information of the static feature information is determined. Target detection is performed on image data from different frames to determine the number of times the target is detected, and confidence feature information of the static feature information is determined based on the number of times.

[0254] Optionally, the first determining module 131 is used to perform some or all of the following steps: Target detection is performed on image data from different frames to determine the center point coordinates of the target bounding box in different frames. Based on the center point coordinates of the target bounding box in different frames and the second duration, the velocity feature information of the dynamic feature information is determined. Based on the velocity feature information and the second duration, the acceleration feature information of the dynamic feature information is determined; The variance of the velocity feature information is calculated to determine the rhythmic feature information of the dynamic feature information.

[0255] Optionally, the switching module 134 is used for: If the first comparison result indicates that the target scene and the current scene are the same, then the current scene remains unchanged; If the first comparison result indicates that the target scene and the current scene are different, then the current scene is switched to the target scene.

[0256] Optionally, after determining whether to switch scenes based on the first comparison result, the switching module 134 is further configured to: Based on the attention value of the target in the current frame, determine the main attention target of the current frame; Based on the target scene, the emotional state information, the main attention target of the current frame, and the historical task information, a target task chain is selected from multiple preset task chains and executed.

[0257] Based on the same inventive concept, this application also provides a robot 140, such as... Figure 14 As shown, it includes at least one processor 141 and a memory 142 connected to at least one processor. In this embodiment, the specific connection medium between the processor 141 and the memory 142 is not limited. Figure 14 Taking the connection between processor 141 and memory 142 via a bus as an example. The bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, Figure 14 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0258] The processor 141 serves as the robot's facial expression and motion coordination control center. It connects to various parts of the robot via various interfaces and wiring, and performs data processing by running or executing instructions stored in the memory 142 and retrieving data stored in the memory 142. Optionally, the processor 141 may include one or more processing units. The processor 141 may integrate an application processor and a modem processor. The application processor primarily handles the operating system, user interface, and applications, while the modem processor primarily handles issuing instructions. It is understood that the modem processor may not be integrated into the processor 141. In some embodiments, the processor 141 and the memory 142 may be implemented on the same chip; in other embodiments, they may be implemented on separate chips.

[0259] Processor 141 can be a general-purpose processor, such as a CPU, digital signal processor, application-specific integrated circuit (ASIC), field-programmable gate array or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the robot expression and motion collaborative control method can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0260] Memory 142, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Memory 142 may include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic storage, magnetic disk, optical disk, etc. Memory 142 can be any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto. In the embodiments of this application, memory 142 may also be a circuit or any other device capable of implementing storage functions for storing program instructions and / or data.

[0261] In this embodiment, the memory 142 stores a computer program, which, when executed by the processor 141, causes the processor 141 to perform the steps of the scene switching method of the robot described above.

[0262] Based on the same inventive concept, embodiments of this application provide a computer-readable storage medium storing a computer program executable by a computer device, which, when run on the computer device, causes the computer device to perform the steps of the above-described scene switching method for the robot.

[0263] Based on the same inventive concept, this application provides a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions, which, when executed by a computer device, cause the computer device to perform the steps of the above-described scene switching method for a robot.

[0264] Those skilled in the art will understand that embodiments of the present invention can be provided as methods or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0265] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer apparatus or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0266] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer device or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0267] These computer program instructions may also be loaded onto a computer device or other programmable data processing equipment to cause a series of operational steps to be performed on the computer device or other programmable equipment to produce a process implemented by the computer device, thereby providing instructions that execute on the computer device or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0268] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0269] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A scene switching method of a robot, characterized by, include: Feature extraction is performed on the multimodal data of the target in different frames to obtain frame feature information, and the target attention value of the target in the current frame is determined based on the frame feature information; The stimulus event of the current frame is determined based on the multimodal data, and the emotional state information of the current frame is determined based on the weight of the stimulus event. The stimulus event is an event that causes a change in the emotional state. The target scene tendency value of the current frame is determined based on the scene tendency value of the previous frame, the target attention value, and the emotional state information. The target scene is determined based on the target scene tendency value, and the target scene is compared with the current scene. Based on the first comparison result, it is determined whether to switch scenes.

2. The method of claim 1, wherein, Determining the target attention value of the target in the current frame based on the frame feature information includes: Based on the frame feature information, the attention value of the target in different frames is determined; The target attention value is determined based on the target's attention values ​​in different frames and the total number of frames.

3. The method of claim 1, wherein, Determining the stimulus event of the current frame based on the multimodal data includes some or all of the following steps: Based on the frame feature information, the burst level information of the current frame is determined. If the burst level information of the current frame exceeds the set level threshold, the stimulus event is generated. Based on the attention values ​​of the target in different frames, the duration of the target as the primary attention target is determined. If the duration of the target as the primary attention target exceeds a first duration threshold, the stimulus event is generated. Emotion recognition is performed on the speech data of multimodal data from different frames. Based on the emotion recognition results, the duration of the speech data corresponding to different emotions is determined. If there is an emotion that exceeds a second duration threshold, the stimulus event is generated. Target detection is performed on image data of multimodal data from different frames. Based on the target detection results, a first duration corresponding to image data that does not include the target is determined. If the first duration exceeds a third duration threshold, the stimulus event is generated. Based on the pressure data from multimodal data of different frames, the duration of touch on different parts of the robot is determined. If the duration exceeds the fourth duration threshold, the stimulus event is generated.

4. The method of claim 3, wherein, Before determining the target scene tendency value of the current frame, the method further includes: Based on the frame feature information, the target burst level information is determined; Determining the target scene tendency value for the current frame based on the scene tendency value of the previous frame, the target attention value, and the emotional state information includes: The target scene tendency value is determined based on the scene tendency value of the previous frame of the current frame, the target attention value, the target suddenness level information, and the emotional state information.

5. The method of claim 4, wherein, The determination of the target burst level information based on the frame feature information includes: Based on the frame feature information and the second duration, the feature change rate of the target in different frames is determined, where the second duration is the duration between two adjacent frames; The burst level information of the target is determined based on the characteristic change rate of the target in different frames.

6. The method of claim 1, wherein, Determining the emotional state information of the current frame based on the weights of the stimulus events includes: If the current time is within the task's execution time period, then determine the initial emotional state information and the termination emotional state information for the execution time period. Based on the duration corresponding to the execution time period, the third duration, the initial emotional state information, and the termination emotional state information, the first emotional state information is determined, and the third duration is determined based on the current time and the earliest time of the execution time period; Based on the weights of the stimulus events, the second emotional state information is determined; The emotional state information is determined based on the first emotional state information and the second emotional state information.

7. The method as described in claim 6, characterized in that, The first emotional state information is determined by the following equation : ; wherein, is initial emotional state information, is termination emotional state information, is an earliest time of an execution time period of the task, is a latest time of an execution time period of the task, is a current time, .

8. The method as described in claim 1, characterized in that, Determining the emotional state information of the current frame based on the weights of the stimulus events includes: If the current time is not within the task's execution time period, the second emotional state information is determined based on the weight of the stimulus event. The second emotional state information is compared with a set threshold to obtain a second comparison result; If the second comparison result indicates that the second emotional state information is less than a set threshold, then it is determined that the robot is in an idle state, and based on the third emotional state information when the robot is first detected to be in an idle state, the set emotional state information, and the fourth duration, the fourth emotional state information is determined, wherein the fourth duration is determined based on the current time and the time when the robot is first detected to be in an idle state; The emotional state information is determined based on the second emotional state information and the fourth emotional state information.

9. The method as described in claim 8, characterized in that, After comparing the second emotional state information with a set threshold to obtain a second comparison result, the method further includes: If the second comparison result indicates that the second emotional state information is greater than or equal to a set threshold, then the second emotional state information is used as the emotional state information of the current frame.

10. The method according to any one of claims 1-5, characterized in that, The frame feature information is determined using the following method: Based on the scene corresponding to different frames, the target and the first correspondence relationship, the baseline feature information of the frame feature information is determined, and the first correspondence relationship is the correspondence relationship between each scene, each target and each baseline feature information. Static feature extraction is performed on the image data of multimodal data from different frames to determine the static feature information of the frame feature information; Dynamic feature extraction is performed on the image data of multimodal data from different frames to determine the dynamic feature information of the frame feature information.

11. The method as described in claim 10, characterized in that, The static feature extraction of image data from multimodal data of different frames to determine the static feature information of the frame feature information includes some or all of the following steps: Target detection is performed on image data from different frames to determine the target bounding boxes of the target in different frames. Based on the center point coordinates of the target bounding boxes in different frames and the center point coordinates of the image data in different frames, the centrality feature information of the static feature information is determined. Based on image data from different frames, a deep learning method is used to determine the target angle of the target in different frames, and based on the target angle of the target in different frames, the orientation feature information of the static feature information is determined, wherein the target angle is the angle between the target normal and the camera's line of sight; Target detection and key point detection are performed on image data of different frames to determine the key points of the target in different frames. Based on the key points of the target in different frames and the depth image data of multimodal data in different frames, the distance feature information of the static feature information is determined. Target detection is performed on image data from different frames to determine the number of times the target is detected, and confidence feature information of the static feature information is determined based on the number of times.

12. The method as described in claim 10, characterized in that, The dynamic feature extraction of image data from multimodal data of different frames to determine the dynamic feature information of the frame feature information includes some or all of the following steps: Target detection is performed on image data from different frames to determine the center point coordinates of the target bounding box in different frames. Based on the center point coordinates of the target bounding box in different frames and the second duration, the velocity feature information of the dynamic feature information is determined. Based on the velocity feature information and the second duration, the acceleration feature information of the dynamic feature information is determined; The variance of the velocity feature information is calculated to determine the rhythmic feature information of the dynamic feature information.

13. The method as described in claim 2, characterized in that, The step of determining whether to switch scenes based on the first comparison result includes: If the first comparison result indicates that the target scene and the current scene are the same, then the current scene remains unchanged; If the first comparison result indicates that the target scene and the current scene are different, then the current scene is switched to the target scene.

14. The method as described in claim 13, characterized in that, After determining whether to switch scenes based on the first comparison result, the method further includes: Based on the attention value of the target in the current frame, determine the main attention target of the current frame; Based on the target scene, the emotional state information, the main attention target of the current frame, and the historical task information, a target task chain is selected from multiple preset task chains and executed.

15. A scene switching device for a robot, characterized in that, include: The first determining module is used to extract features from the multimodal data of the target in different frames to obtain frame feature information, and to determine the target attention value of the target in the current frame based on the frame feature information. The second determining module is used to determine the stimulus event of the current frame based on the multimodal data, and to determine the emotional state information of the current frame based on the weight of the stimulus event, wherein the stimulus event is an event that causes a change in the emotional state. The third determining module is used to determine the target scene tendency value of the current frame based on the scene tendency value of the previous frame of the current frame, the target attention value, and the emotional state information. The switching module is used to determine the target scene based on the target scene tendency value, compare the target scene with the current scene, and determine whether to switch scenes based on the first comparison result.

16. A robot, characterized in that, It includes at least one processor and at least one memory, wherein the memory stores a computer program that, when executed by the processor, causes the processor to perform the scene switching method for the robot as described in any one of claims 1 to 14.