Information interaction method and device, robot and storage medium

By collecting and analyzing multimodal information, determining the dynamic context constraint set, and combining it with a large language model to optimize robot interaction, the problem of companion robots lacking natural and emotional experience is solved, and a more natural and emotional interaction process is achieved.

CN121936497APending Publication Date: 2026-04-28UBTECH ROBOTICS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
UBTECH ROBOTICS CORP LTD
Filing Date
2026-02-26
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing companion robots cannot simulate a natural and warm companionship experience, lack situational adaptability, cannot perceive non-verbal information such as facial expressions and body movements, their interactions lack warmth, and their output content lacks immersion.

Method used

By collecting multimodal information from users and extracting features, a contextual feature vector is obtained. Based on the multimodal information and the contextual feature vector, a dynamic contextual constraint set is determined, and the robot is controlled to output interactive information. Combined with a large language model, content generation is optimized to achieve emotional interaction.

Benefits of technology

It achieves naturalness and emotionalization in the robot interaction process, and can adjust the output content according to the user's context and emotions, thereby enhancing the user's immersion and participation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121936497A_ABST
    Figure CN121936497A_ABST
Patent Text Reader

Abstract

The invention provides an information interaction method and device, a robot and a storage medium, and belongs to the technical field of robotics.The method comprises the steps that in response to an information interaction instruction, multi-modal information of a user is collected; feature extraction is carried out on the multi-modal information to obtain a situation feature vector, and the situation feature vector is used for representing a vector of a current interaction situation; based on the situation feature vector and the multi-modal information, a dynamic situation constraint set is determined, and the dynamic situation constraint set is used for controlling interaction information output by the robot; and controlling the robot to output interaction information based on the dynamic situation constraint set. In this way, the current scene is comprehensively perceived through multi-modal information, a traditional mechanical interaction mode is broken, and the information interaction process is more natural and emotional.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of robotics technology, and in particular relates to an information interaction method, device, robot, and storage medium. Background Technology

[0002] With the development of intelligent robots, companion robots are gradually being applied in scenarios such as homes and education. For example, companion robots such as children's story machines are becoming increasingly common in applications that accompany young users.

[0003] Current companion robots typically use speech recognition technology for information interaction. They trigger interaction with wake words, convert user voice commands into text information, and then search for and play audio stories that match the text information through cloud retrieval.

[0004] However, the information interaction scenarios of the aforementioned companion robots rely on keyword detection and matching and a fixed audio library, which can only achieve basic voice information interaction functions and cannot simulate a natural and warm companionship experience. Summary of the Invention

[0005] The purpose of this application is to provide an information interaction method, device, robot, and storage medium, which aims to solve the problem that current robots cannot simulate a natural and warm companionship experience.

[0006] A first aspect of this application provides an information interaction method, the method comprising:

[0007] Responding to information interaction commands, it collects multimodal information from users; Feature extraction is performed on the multimodal information to obtain a context feature vector, which is used to represent the current interaction context. Based on the context feature vector and the multimodal information, a dynamic context constraint set is determined, which is used to control the interactive information output by the robot. Based on the dynamic situation constraint set, the robot is controlled to output interactive information.

[0008] In some embodiments, determining the dynamic situation constraint set based on the situation feature vector and the multimodal information includes: Based on the context feature vector and the preset time rules, scene category recognition is performed to obtain the target interaction scene; Based on the multimodal information, user emotion recognition is performed to obtain the user's emotion dimension information; Based on the multimodal information, the current interactive content is identified to obtain real-time interactive information; The target interaction scenario, the emotional dimension information, and the real-time interaction information are combined to form the dynamic situation constraint set.

[0009] In some embodiments, controlling the robot to output interactive information based on the dynamic context constraint set includes: The dynamic context constraint set is encoded into a control signal for a large language reasoning model, and the control signal is used to guide the output of the large language model. Based on the large language model, the interactive information to be output is determined under the control signals corresponding to the dynamic context constraint set; Control the robot to output the interactive information.

[0010] In some embodiments, determining the interaction information to be output based on the large language model and under the control signal corresponding to the dynamic context constraint set includes: Based on the control signal, determine the user's target emotional dimension information; Based on the difference between the target emotion dimension information and the user's emotion dimension information, the robot's interaction information is determined, which is used to guide the user's emotion dimension information to the target emotion dimension information.

[0011] In some embodiments, determining the robot's interaction information based on the difference between the target emotion dimension information and the user's emotion dimension information includes: Based on the difference between the target emotion dimension information and the user's emotion dimension information, adjust the word type and quantity in the content information; and... Based on the difference between the target emotion dimension information and the user's emotion dimension information, the robot's expression parameters are determined, and the expression parameters are used to characterize the output method of the robot's content information.

[0012] In some embodiments, the multimodal information includes audio information, visual information, and environmental information, and the step of extracting features from the multimodal information to obtain a contextual feature vector includes: Feature extraction is performed on the audio information to obtain the user's voiceprint features, volume features, pitch features, speech rate features, and audio emotion features; Feature extraction is performed on the visual information to obtain the user's body movement features and facial expression features; Feature extraction is performed on the environmental information to obtain the light intensity features, color temperature features, and noise level features of the user's environment; The context feature vector is obtained by weighted fusing the voiceprint features, volume features, pitch features, speech rate features, audio emotion features, body movement features, facial expression features, light intensity features, color temperature features, and noise level features.

[0013] In some embodiments, the method further includes: In response to receiving a user-triggered interaction request, determine that an information interaction instruction has been received; or, In response to a pre-defined interaction node in the information interaction scenario, it is determined that an information interaction instruction has been received.

[0014] A second aspect of this application provides an information interaction device, the device comprising: The information acquisition unit is used to collect multimodal information from users in response to information interaction commands; The feature extraction unit is used to extract features from the multimodal information to obtain a context feature vector, which is used to represent the current interaction context. The determining unit is used to determine a dynamic situation constraint set based on the situation feature vector and the multimodal information, wherein the dynamic situation constraint set is used to control the interactive information output by the robot. The control unit is used to control the robot to output interactive information based on the dynamic situation constraint set.

[0015] In some embodiments, the determining unit is configured to: identify a scene category based on the context feature vector and a preset time rule to obtain a target interaction scene; identify user emotions based on the multimodal information to obtain user emotion dimension information; identify the current interaction content based on the multimodal information to obtain real-time interaction information; and combine the target interaction scene, the emotion dimension information, and the real-time interaction information to form the dynamic context constraint set.

[0016] In some embodiments, the control unit is configured to encode the dynamic context constraint set into a control signal for a large language reasoning model, the control signal being used to guide the output of the large language model; based on the large language model, under the control signal corresponding to the dynamic context constraint set, determine the interactive information to be output; and control the robot to output the interactive information.

[0017] In some embodiments, the control unit is configured to determine the user's target emotional dimension information based on the control signal; and to determine the robot's interaction information based on the difference between the target emotional dimension information and the user's emotional dimension information, wherein the interaction information is used to guide the user's emotional dimension information to the target emotional dimension information.

[0018] In some embodiments, the control unit is configured to adjust the word type and quantity in the content information based on the difference between the target emotion dimension information and the user's emotion dimension information; and to determine the robot's expression parameters based on the difference between the target emotion dimension information and the user's emotion dimension information, the expression parameters being used to characterize the output method of the robot's content information.

[0019] In some embodiments, the multimodal information includes audio information, visual information, and environmental information. The feature extraction unit is used to extract features from the audio information to obtain the user's voiceprint features, volume features, pitch features, speech rate features, and audio emotion features; to extract features from the visual information to obtain the user's body movement features and facial expression features; to extract features from the environmental information to obtain the light intensity features, color temperature features, and noise level features of the user's environment; and to perform weighted fusion of the voiceprint features, volume features, pitch features, speech rate features, audio emotion features, body movement features, facial expression features, light intensity features, color temperature features, and noise level features to obtain the context feature vector.

[0020] In some embodiments, the apparatus further includes: The receiving unit is configured to, in response to receiving a user-triggered interaction request, determine that an information interaction instruction has been received; or, The receiving unit is used to respond to a preset interaction node in the information interaction scenario and determine that an information interaction instruction has been received.

[0021] A third aspect of this application provides a robot including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the information interaction method described above.

[0022] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the information interaction method described above.

[0023] A fifth aspect of this application provides a computer program product that, when run on an electronic device, causes the electronic device to perform the information interaction method described above.

[0024] The beneficial effects of the embodiments of the present invention compared with the prior art are as follows: In this embodiment, by collecting the user's multimodal information and extracting features from the multimodal information, a contextual feature vector is obtained to characterize the current interaction context. Based on the multimodal information and the contextual feature vector, the current dynamic contextual constraint set is determined. Based on the dynamic contextual constraint set, the robot is controlled to output interactive information. In this way, the current context is perceived comprehensively through multimodal information, breaking the traditional mechanical interaction mode and making the information interaction process more natural and emotional. Attached Figure Description

[0025] Figure 1 A schematic diagram of the system architecture of a robot provided in an exemplary embodiment of this application is shown; Figure 2 A flowchart illustrating an exemplary embodiment of an information interaction method is shown. Figure 3 A flowchart illustrating a method for determining a set of dynamic context constraints provided by an exemplary embodiment is shown. Figure 4 A flowchart illustrating an exemplary embodiment of an information interaction method is shown. Figure 5 A schematic diagram of a method for determining information interaction of a robot provided by an exemplary embodiment is shown; Figure 6 A schematic diagram of the structure of an information interaction device provided in this application is shown; Figure 7 A schematic diagram of the structure of a robot provided in this application is shown. Detailed Implementation

[0026] To make the technical problems, technical solutions, and beneficial effects to be solved by this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the scope of this application.

[0027] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0028] With the development of intelligent robots, companion robots are gradually being applied in scenarios such as homes and education. For example, companion robots such as children's story machines are becoming increasingly common in applications that accompany young users.

[0029] Current companion robots typically use speech recognition technology for information interaction. Interaction is triggered by a wake word, and the user's voice commands are converted into text information. Then, the system searches the cloud for and plays audio stories that match the text. The robot only supports the "voice wake-up - keyword search - mechanical playback" chain.

[0030] However, the information interaction scenarios of the aforementioned companion robots rely on keyword detection and matching and a fixed audio library, which can only achieve basic voice information interaction functions. They cannot perceive non-verbal information such as facial expressions and body movements of users, and cannot simulate a natural and warm companionship experience, resulting in a lack of warmth in the interaction.

[0031] Furthermore, current companion robots cannot differentiate between usage scenarios. For example, in a "bedtime" scenario, if the robot plays upbeat, stimulating audio, it may actually disrupt the user's sleep; or in an "outdoor" scenario, overly bland content will fail to capture attention. Therefore, the robots lack situational adaptability.

[0032] In other embodiments, the content output by the robot is mostly pre-recorded or stored information, which users can only passively listen to and cannot participate in the development of the story, making it difficult to stimulate users' imagination and creative participation. In addition, the inability to convert visual information into story elements in real time results in insufficient immersion.

[0033] To improve the naturalness of robot interaction and optimize the companionship experience, this application provides an information interaction method, device, robot, and storage medium. By collecting multimodal information from the user and extracting features from this information, a contextual feature vector representing the current interaction situation is obtained. Based on this multimodal information and contextual feature vector, a dynamic contextual constraint set is determined. The robot is then controlled to output interactive information based on this dynamic contextual constraint set. This comprehensive perception of the current situation through multimodal information breaks away from the traditional mechanical interaction mode, making the information interaction process more natural and emotional.

[0034] The present application will now be described with reference to specific embodiments. See below. Figure 1 The illustration shows a system architecture of a robot provided by an exemplary embodiment of this application, which includes: a multimodal perception and fusion module, a scene reasoning and state judgment module, a content generation and adaptive adjustment module, and a multimodal output and interaction module.

[0035] The multimodal perception and fusion module is used to identify multimodal information such as speech, visual, and environmental information, and to fuse the identified feature information. In some embodiments, the multimodal perception and fusion module deploys a cross-modal attention network (CMAN), which performs weighted fusion of multimodal data and outputs a unified contextual feature vector.

[0036] It should be noted that the robot also includes a data acquisition system, which comprises multiple information acquisition modules for collecting the user's multimodal information. For example, the data acquisition system includes a microphone module for collecting the user's audio information; a camera for collecting the user's visual information; and environmental sensors for collecting environmental information about the user's surroundings. This data acquisition system communicates with the multimodal perception and fusion module, transmitting the collected multimodal information to it.

[0037] The scenario reasoning and state judgment module is used for contextual reasoning through a two-layer reasoning mechanism. This two-layer mechanism includes macro-level scenario recognition and micro-level state judgment. Macro-level scenario recognition identifies the target scenario based on contextual feature vectors and preset time rules. In some embodiments, the scenario reasoning and state judgment module deploys a neural network for scenario recognition; for example, this neural network can be a Hidden Markov Model (HMM), a Recurrent Neural Network (RNN), or a Long Short-Term Memory (LSTM) network. The micro-level state judgment is used to determine the user's emotional dimension information in real time, including arousal (A) and valence (V), with the output data in the form of [A, V] coordinate values. The arousal level represents the physiological activation or intensity of an emotion, describing a continuous spectrum from "low activation / sleep state" to "high activation / excitement state"; the affective valence represents the positive or negative aspect or pleasantness of an emotion, describing a continuous spectrum from extreme "negative / unpleasant" to extreme "positive / pleasant". Correspondingly, the [A, V] coordinate values ​​represent the position of different emotions in the emotion coordinate system. For example, the upper right quadrant (high valence, high arousal) represents emotions such as excitement, ecstasy, and exhilaration; the lower right quadrant (high valence, low arousal) represents emotions such as tranquility, satisfaction, and relaxation; the upper left quadrant (low valence, high arousal) represents emotions such as anger, anxiety, fear, and panic; the lower left quadrant (low valence, low arousal) represents emotions such as sadness, depression, fatigue, and boredom; and the area near the central origin represents neutral and apathetic emotions.

[0038] The content generation and adaptive adjustment module is used to generate the content information to be output based on information such as the current context feature vector. This module includes a large language model as its core. This large language model (LLM) is a constrained large language model optimized for the robot's application scenarios. For example, when the robot is a child companion robot, the large language model can be a constrained large language model optimized based on children's cognitive development and educational principles.

[0039] This multimodal output and interaction module controls the robot to output interactive information, triggers information interaction commands at interaction nodes, and issues prompts. Furthermore, it determines the real-time story variable input (SVI) based on the user's multimodal information and uses the SVI as the core input variable for the next round of LLM generation to drive the causal process of the interaction.

[0040] In summary, by collecting multimodal information from users and extracting features from this information, a contextual feature vector representing the current interaction context is obtained. Based on this multimodal information and contextual feature vector, the current dynamic contextual constraint set is determined. Based on this dynamic contextual constraint set, the robot is controlled to output interactive information. In this way, by comprehensively perceiving the current context through multimodal information, the traditional mechanical interaction mode is broken, making the information interaction process more natural and emotional.

[0041] The information interaction method provided in this application will be explained below with reference to the specific implementation process. See [link to relevant documentation]. Figure 2 The diagram illustrates a flowchart of an exemplary method for information interaction provided by a robot, as an example and not a limitation.

[0042] S201, in response to information interaction commands, the robot collects multimodal information from the user.

[0043] This information interaction instruction is used to trigger interaction between the robot and the user. In some embodiments, the information interaction instruction is a user-triggered instruction; accordingly, in response to receiving a user-triggered request to initiate interaction, the robot determines that it has received the information interaction instruction. For example, the user can trigger the information interaction instruction via voice. The robot acquires the user's voice information, and if it detects keywords through the voice information, it determines that it has received the information interaction instruction.

[0044] In other embodiments, the robot sets interactive nodes within the interactive content, and the information interaction instructions can be interactive instructions triggered based on these interactive nodes. Accordingly, in response to preset interactive nodes in the information interaction scenario, the robot determines that it has received an information interaction instruction. For example, if the interactive content is a story, the robot can insert interactive nodes during the story's development, or the developers can insert interactive nodes within the story. When the story reaches a preset interactive node, the robot determines that it has received an information interaction instruction. It should be noted that when the story reaches a preset interactive node, the robot can also output prompts to encourage the user to interact. For example, the robot can output audio prompts such as, "The little protagonist is very sad now, please gently stroke it."

[0045] This multimodal information includes various types of information used to characterize the user's context. For example, it includes the user's audio information, visual information, and environmental information. Accordingly, in this step, the robot can collect audio information emitted by the user through a microphone module; the robot can also collect visual information through a camera, including the user's body movements and facial expressions; and the robot can also collect environmental information about the user's environment through environmental sensors, including temperature, humidity, light intensity, and noise levels.

[0046] S202, the robot extracts features from the multimodal information to obtain a context feature vector, which is used to represent the current interaction context.

[0047] The robot extracts features from different types of information within the multimodal information, and then fuses these features to obtain a contextual feature vector. For example, when the multimodal information includes audio, visual, and environmental information, the robot extracts features from the audio information to obtain the user's voiceprint, volume, pitch, speech rate, and emotional characteristics; it extracts features from the visual information to obtain the user's body movements and facial expressions; and it extracts features from the environmental information to obtain the light intensity, color temperature, and noise level of the user's environment. The robot then performs a weighted fusion of these features to obtain the contextual feature vector. The robot can use CMAN to perform weighted fusion of multimodal data and output a unified contextual feature vector.

[0048] It should be noted that, to ensure the accuracy of the contextual feature vector, the robot aligns various pieces of information in the multimodal data before this step. Accordingly, the robot aligns the various pieces of information based on the acquisition timestamp of each piece of information in the multimodal data, and then performs feature extraction and feature fusion on the aligned multimodal data to obtain the contextual feature vector.

[0049] S203, the robot determines a dynamic situation constraint set based on the situation feature vector and the multimodal information. This dynamic situation constraint set is used to control the interactive information output by the robot.

[0050] This dynamic context constraint set is used to represent the current interaction scenario information, thereby controlling the interactive information output by the robot. In some embodiments, the dynamic context constraint set includes a target interaction scenario, user emotion dimension information, and real-time interaction information. The robot determines the target interaction scenario, user emotion dimension information, and real-time interaction information respectively, and then assembles these three elements into the dynamic context constraint set. See also Figure 3 This illustrates a flowchart of a method for determining a set of dynamic situational constraints provided by an exemplary embodiment. Figure 3 As shown, this method can be implemented through the following steps S2031-S2034, including: S2031, the robot performs scene category recognition based on the context feature vector and preset time rules to obtain the target interaction scene.

[0051] The target interaction scenario refers to the interaction scenario that matches the current context feature vector and the preset time rules. The robot can support multiple interaction scenarios. These multiple interaction scenarios can be those set by the robot developers or those customized after the robot leaves the factory; in this embodiment, no specific limitation is made. For example, the interaction scenario can be a bedtime scenario, a play scenario, an outdoor scenario, and a learning scenario, etc.

[0052] This time rule is a pre-defined correspondence between interactive scenarios and time. For example, 10 PM is set as the time before bedtime, and 7 PM to 9 PM is set as the time for learning. This pre-defined time rule can be an interactive scenario set by the robot developers or a custom interactive scenario defined after the robot leaves the factory; in this embodiment, no specific limitation is made.

[0053] It's important to note that this target interaction scenario can also be one determined by the user's direct interaction with the robot. For example, the user can select the target interaction scenario through the robot's human-computer interface. Alternatively, the user can select the target interaction scenario through voice interaction with the robot. For instance, the user can say to the robot, "I want to start learning," and the robot, through semantic analysis, determines that the user wishes to enter a learning state, thus defining the target interaction scenario as the learning scenario.

[0054] In some embodiments, the robot identifies the target interaction scenario using an HMM model or an RNN / LSTM model.

[0055] S2032, the robot performs user emotion recognition based on this multimodal information to obtain the user's emotional dimension information.

[0056] In this embodiment, the emotional dimension information is represented by arousal (A) - valence (V), and the emotional dimension information is represented by the [A, V] coordinate values. The robot can identify the user's emotional dimension information by combining visual and audio information from multimodal information. That is, the robot analyzes the user's micro-expression information through visual information, and determines the user's emotional dimension information by combining audio information and behavioral actions.

[0057] S2033, the robot identifies the current interactive content based on this multimodal information and obtains real-time interactive information.

[0058] The robot captures multimodal information, including body movements, and combines this information with the current interaction content to determine the other content that should be output in the preset interaction flow in order, as real-time interaction information. For example, if the current interaction content is "The little protagonist is very sad now, please gently stroke it," and the robot recognizes the user's body movement as a "stroking" action, it determines that the real-time interaction information has been completed according to the current interaction requirements, and then outputs preset encouraging content.

[0059] S2034, the robot combines the target interaction scenario, the emotional dimension information, and the real-time interaction information into the dynamic situation constraint set.

[0060] In this step, the robot combines multiple contextual constraint information into a dynamic contextual constraint set according to a preset data format. For example, this dynamic contextual constraint set Pconstraint={target scene, [A, V], real-time interaction information SVI}.

[0061] In this implementation, by combining environmental information, user status, and real-time user interaction information into a dynamic context constraint set, the robot can determine the interaction information by combining multiple information when outputting interaction information, so that the interaction information can better match the user's current interaction context.

[0062] S204, the robot controls the output of interactive information based on this dynamic situation constraint set.

[0063] In this embodiment, the robot generates the interactive information through a content generation and adaptive adjustment module, which includes a large language model as the main model. This large language model (LLM) is a constrained large language model optimized for the robot's application scenario. Accordingly, the robot uses the dynamic situational constraint set as input to the large language model, and then outputs the interactive information through the large language model.

[0064] S2043, the robot controls the robot to output the interactive information.

[0065] In the embodiments of this application, the robot can output the interactive information through multiple modalities. Accordingly, the robot can determine the corresponding expression parameters based on the output of different modalities, and the robot outputs the corresponding interactive information based on the expression parameters corresponding to at least one modality.

[0066] In some embodiments, when the interaction information is audio information, the expression parameter may include speech rate and fundamental frequency. Accordingly, when the interaction information is audio information, the robot determines the speech rate and fundamental frequency of the interaction information. For example, when the dynamic context constraint set indicates that the user prefers "concerned" voice interaction information, the robot determines the expression parameter as speech rate. 10%, base frequency +5%; when the dynamic context constraint set indicates that the user prefers "humorous" voice interaction information, the robot determines the expression parameters as speech rate +8% and base frequency +12%.

[0067] In some embodiments, the interactive information may also be visual information. This visual information may be multimedia information, or it may be robot gesture information, etc. Accordingly, the expression parameters include display hue and / or robot control parameters. Accordingly, when the interactive information is visual information, the robot determines the display hue and / or robot control parameters of the interactive information based on the dynamic context constraint set.

[0068] For example, when the interactive information is multimedia information, the expression parameter can be a multi-segment LED display on the head-mounted circular screen, showing a breathing animation in cool or warm colors that matches the dynamic context constraint set. When the interactive information is robot gesture information, the expression parameter can be a robot shoulder degree of freedom (DOF) of 2 and an elbow DOF of 1. In some embodiments, the robot can pre-store preset gesture templates corresponding to different emotions. Accordingly, in this step, the robot can determine the expression parameters of the corresponding gesture template from the preset gesture templates based on the dynamic context constraint set.

[0069] In this embodiment, by collecting the user's multimodal information and extracting features from the multimodal information, a contextual feature vector is obtained to characterize the current interaction context. Based on the multimodal information and the contextual feature vector, the current dynamic contextual constraint set is determined. Based on the dynamic contextual constraint set, the robot is controlled to output interactive information. In this way, the current context is perceived comprehensively through multimodal information, breaking the traditional mechanical interaction mode and making the information interaction process more natural and emotional.

[0070] See Figure 4 The diagram illustrates a flowchart of an exemplary method for information interaction provided by a robot, as an example and not a limitation.

[0071] S401, in response to information interaction commands, the robot collects multimodal information from the user.

[0072] This step is based on the same principle as step S201, and will not be repeated here.

[0073] S402, the robot extracts features from the multimodal information to obtain a context feature vector, which is used to represent the current interaction context.

[0074] This step is based on the same principle as step S202, and will not be repeated here.

[0075] S403, the robot determines a dynamic situation constraint set based on the situation feature vector and the multimodal information. This dynamic situation constraint set is used to control the interactive information output by the robot.

[0076] This step is based on the same principle as step S203, and will not be repeated here.

[0077] S404, the robot encodes the dynamic situation constraint set into control signals for the large language reasoning model, which are used to guide the output of the large language model.

[0078] This control signal guides the Large Language Model (LLM) to output the result of matching the dynamic context constraint set. In some embodiments, the control signal is a prefix prompt for the LLM, and the robot encodes the dynamic context constraint set as a prefix prompt. In some embodiments, the dynamic context constraint set can be directly inserted as a prefix prompt into the LLM input, for example, using "Please write a calm, warm, and slow bedtime story paragraph" as the LLM input. In some embodiments, the control signal is an internal attention bias for the LLM, and the robot encodes the dynamic context constraint set as an internal attention bias. For example, the dynamic context constraint set can be used as an attention bias to fine-tune the generation probability of certain words within the LLM.

[0079] The large language model outputs the result corresponding to the dynamic context constraint set, namely the Story Segment. _t =LLM(Theme, History, Pconstraint_t), where Theme represents the theme of the interactive information to be generated, History identifies the content information generated in the past, and Pconstraint_t represents the set of dynamic context constraints at time t.

[0080] S405, based on this large language model, the robot determines the interactive information to be output under the control signals corresponding to the dynamic situation constraint set.

[0081] In this embodiment, the robot determines the target emotion dimension information matching the dynamic situation constraint set based on the control signal. Then, based on the difference between the target emotion dimension information and the user's current emotion position information, the robot adjusts the interaction information so that the user can reach the emotion corresponding to the target emotion dimension information under the guidance of the robot's output interaction information. Accordingly, see... Figure 5 This illustrates a flowchart of a process for determining robot interaction information, provided by an exemplary embodiment. Figure 5 As shown, this process can be implemented through the following steps S4051-S4052, including: S4051, the robot determines the user's target emotional dimension information based on this control signal.

[0082] The robot sets target emotional information based on the target interaction scenario. This target emotional information includes the target emotional range corresponding to the target interaction scenario. For example, in the bedtime scenario, the target emotional dimension is low A (low arousal) and high V (emotional valence). The robot can store the value range of [A, V] corresponding to the bedtime scenario.

[0083] S4052, the robot determines the robot's interaction information based on the difference between the target emotional dimension information and the user's emotional dimension information. This interaction information is used to guide the user's emotional dimension information to the target emotional dimension information.

[0084] The robot determines its interaction information based on the difference between the target emotion dimension information and the user's current emotion dimension information. This interaction information includes output content information and expression parameters. Accordingly, the robot adjusts the word types and quantities in the content information based on the difference between the target emotion dimension information and the user's emotion dimension information; and, based on the difference between the target emotion dimension information and the user's emotion dimension information, it determines the robot's expression parameters, which characterize the way the robot outputs its content information.

[0085] For example, if the current user's emotional dimension information indicates that user A is too high (i.e., too excited), reduce strong action verbs (such as "run," "jump," or "explode") and increase mild verbs (such as "stroll" or "lie down gently"), while also slowing down the speech rate. If the current user's emotional dimension information indicates that user V is too low (i.e., unhappy), increase positive adjectives (such as "warm," "gentle," or "comfortable") and reduce negative words, while also increasing the warmth and gentleness of the tone.

[0086] S406, the robot controls the robot to output the interactive information.

[0087] This step is based on the same principle as step S204, and will not be repeated here.

[0088] In this embodiment, by collecting the user's multimodal information and extracting features from the multimodal information, a contextual feature vector is obtained to characterize the current interaction context. Based on the multimodal information and the contextual feature vector, the current dynamic contextual constraint set is determined. Based on the dynamic contextual constraint set, the robot is controlled to output interactive information. In this way, the current context is perceived comprehensively through multimodal information, breaking the traditional mechanical interaction mode and making the information interaction process more natural and emotional.

[0089] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0090] See Figure 6 It shows a schematic diagram of the structure of an information interaction device provided in this application, including various units used to perform the various steps in the above embodiments, see [link to diagram]. Figure 6 The information interaction device includes: The information acquisition unit 601 is used to collect the user's multimodal information in response to information interaction commands; The feature extraction unit 602 is used to extract features from the multimodal information to obtain a context feature vector, which is used to represent the current interaction context. The determining unit 603 is used to determine a dynamic situation constraint set based on the situation feature vector and the multimodal information. The dynamic situation constraint set is used to control the interactive information output by the robot. The control unit 604 is used to control the robot to output interactive information based on the dynamic situation constraint set.

[0091] In some embodiments, the determining unit 603 is configured to perform scene category recognition based on the context feature vector and preset time rules to obtain a target interaction scene; perform user emotion recognition based on the multimodal information to obtain user emotion dimension information; identify the current interaction content based on the multimodal information to obtain real-time interaction information; and form the target interaction scene, the emotion dimension information and the real-time interaction information into the dynamic context constraint set.

[0092] In some embodiments, the control unit 604 is configured to encode the dynamic context constraint set into a control signal for a large language reasoning model, the control signal being used to guide the output of the large language model; based on the large language model, under the control signal corresponding to the dynamic context constraint set, determine the interactive information to be output; and control the robot to output the interactive information.

[0093] In some embodiments, the control unit 604 is configured to determine the user's target emotional dimension information based on the control signal; and to determine the robot's interaction information based on the difference between the target emotional dimension information and the user's emotional dimension information, the interaction information being used to guide the user's emotional dimension information to the target emotional dimension information.

[0094] In some embodiments, the control unit 604 is configured to adjust the word type and quantity in the content information based on the difference between the target emotion dimension information and the user's emotion dimension information; and to determine the robot's expression parameters based on the difference between the target emotion dimension information and the user's emotion dimension information, the expression parameters being used to characterize the output method of the robot's content information.

[0095] In some embodiments, the multimodal information includes audio information, visual information, and environmental information. The feature extraction unit 602 is used to extract features from the audio information to obtain the user's voiceprint features, volume features, pitch features, speech rate features, and audio emotion features; to extract features from the visual information to obtain the user's body movement features and facial expression features; to extract features from the environmental information to obtain the light intensity features, color temperature features, and noise level features of the user's environment; and to perform weighted fusion of the voiceprint features, volume features, pitch features, speech rate features, audio emotion features, body movement features, facial expression features, light intensity features, color temperature features, and noise level features to obtain the context feature vector.

[0096] In some embodiments, the device further includes: The receiving unit is configured to, in response to receiving a user-triggered interaction request, determine that an information interaction instruction has been received; or, The receiving unit is used to respond to a preset interaction node in the information interaction scenario and determine that an information interaction instruction has been received.

[0097] In this embodiment, by collecting the user's multimodal information and extracting features from the multimodal information, a contextual feature vector is obtained to characterize the current interaction context. Based on the multimodal information and the contextual feature vector, the current dynamic contextual constraint set is determined. Based on the dynamic contextual constraint set, the robot is controlled to output interactive information. In this way, the current context is perceived comprehensively through multimodal information, breaking the traditional mechanical interaction mode and making the information interaction process more natural and emotional.

[0098] Figure 7 This is a schematic diagram of a robot provided in an exemplary embodiment of this application. Figure 7 As shown, the robot 7 in this embodiment includes a processor 70, a memory 71, and a computer program 72 stored in the memory 71 and executable on the processor 70, such as an information interaction program. When the processor 70 executes the computer program 72, it implements the steps described in the various information interaction method embodiments above, for example... Figure 2 The steps S201 to S204 are shown. Alternatively, when the processor 70 executes the computer program 72, it implements the functions of each unit in the above-described device embodiments, for example... Figure 6 The functions of units 601 to 604 are shown.

[0099] For example, the computer program 72 can be divided into one or more units, which are stored in the memory 71 and executed by the processor 70 to complete this application. The one or more units can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program 72 in the robot 7. For example, the computer program 72 can be divided into an information acquisition unit 601, a feature extraction unit 602, a determination unit 603, and a control unit 604, with the specific functions of each module as follows: The information acquisition unit 601 is used to collect the user's multimodal information in response to information interaction commands; The feature extraction unit 602 is used to extract features from the multimodal information to obtain a context feature vector, which is used to represent the current interaction context. The determining unit 603 is used to determine a dynamic situation constraint set based on the situation feature vector and the multimodal information. The dynamic situation constraint set is used to control the interactive information output by the robot. The control unit 604 is used to control the robot to output interactive information based on the dynamic situation constraint set.

[0100] The robot 7 can be any robot with information interaction capabilities. The robot 7 may include, but is not limited to, a processor 70 and a memory 71. Those skilled in the art will understand that... Figure 7 This is merely an example of robot 7 and does not constitute a limitation on robot 7. It may include more or fewer parts than shown, or combine certain parts, or different parts. For example, robot 7 may also include input / output devices, network access devices, buses, etc.

[0101] The processor 70 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0102] The memory 71 can be an internal storage unit of the robot 7, such as a hard drive or memory. The memory 71 can also be an external storage device of the robot 7, such as a plug-in hard drive, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the robot 7. Furthermore, the memory 71 can include both internal and external storage units of the robot 7. The memory 71 is used to store the computer program and other programs and data required by the terminal device. The memory 71 can also be used to temporarily store data that has been output or will be output.

[0103] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0104] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0105] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0106] In the embodiments provided in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0107] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0108] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0109] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.

[0110] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the above method embodiments.

[0111] This application also provides a computer program product that, when run on a mobile terminal, enables the mobile terminal to implement the steps described in the various method embodiments above.

[0112] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. An information exchange method, characterized in that, The method includes: Responding to information interaction commands, it collects multimodal information from users; Feature extraction is performed on the multimodal information to obtain a context feature vector, which is used to represent the current interaction context. Based on the context feature vector and the multimodal information, a dynamic context constraint set is determined, which is used to control the interactive information output by the robot. Based on the dynamic situation constraint set, the robot is controlled to output interactive information.

2. The method as described in claim 1, characterized in that, The step of determining the dynamic situation constraint set based on the situation feature vector and the multimodal information includes: Based on the context feature vector and the preset time rules, scene category recognition is performed to obtain the target interaction scene; Based on the multimodal information, user emotion recognition is performed to obtain the user's emotion dimension information; Based on the multimodal information, the current interactive content is identified to obtain real-time interactive information; The target interaction scenario, the emotional dimension information, and the real-time interaction information are combined to form the dynamic situation constraint set.

3. The method as described in claim 1, characterized in that, The step of controlling the robot to output interactive information based on the dynamic situation constraint set includes: The dynamic context constraint set is encoded into a control signal for a large language reasoning model, and the control signal is used to guide the output of the large language model. Based on the large language model, the interactive information to be output is determined under the control signals corresponding to the dynamic context constraint set; Control the robot to output the interactive information.

4. The method as described in claim 3, characterized in that, Based on the large language model, and under the control signals corresponding to the dynamic context constraint set, the interaction information to be output is determined, including: Based on the control signal, determine the user's target emotional dimension information; Based on the difference between the target emotion dimension information and the user's emotion dimension information, the robot's interaction information is determined, which is used to guide the user's emotion dimension information to the target emotion dimension information.

5. The method as described in claim 4, characterized in that, The step of determining the robot's interaction information based on the difference between the target emotion dimension information and the user's emotion dimension information includes: Based on the difference between the target emotion dimension information and the user's emotion dimension information, adjust the word type and quantity in the content information; and... Based on the difference between the target emotion dimension information and the user's emotion dimension information, the robot's expression parameters are determined, and the expression parameters are used to characterize the output method of the robot's content information.

6. The method as described in claim 1, characterized in that, The multimodal information includes audio information, visual information, and environmental information. The step of extracting features from the multimodal information to obtain a contextual feature vector includes: Feature extraction is performed on the audio information to obtain the user's voiceprint features, volume features, pitch features, speech rate features, and audio emotion features; Feature extraction is performed on the visual information to obtain the user's body movement features and facial expression features; Feature extraction is performed on the environmental information to obtain the light intensity features, color temperature features, and noise level features of the user's environment; The context feature vector is obtained by weighted fusing the voiceprint features, volume features, pitch features, speech rate features, audio emotion features, body movement features, facial expression features, light intensity features, color temperature features, and noise level features.

7. The method according to any one of claims 1-6, characterized in that, The method further includes: In response to receiving a user-triggered interaction request, determine that an information interaction instruction has been received; or, In response to a pre-defined interaction node in the information interaction scenario, it is determined that an information interaction instruction has been received.

8. An information interaction device, characterized in that, The device includes: The information acquisition unit is used to collect multimodal information from users in response to information interaction commands; The feature extraction unit is used to extract features from the multimodal information to obtain a context feature vector, which is used to represent the current interaction context. The determining unit is used to determine the dynamic situation constraint set based on the situation feature vector and the multimodal information; The control unit is used to control the robot to output interactive information based on the dynamic situation constraint set.

9. A robot comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the information interaction method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the information interaction method as described in any one of claims 1 to 7.