Robot interaction processing method, device and equipment, computer readable storage medium and computer program product
By configuring the robot's target personality and utilizing multimodal interaction methods, and combining interaction feedback data to update the personality in real time, the problems of lack of emotion and mechanical nature in robot interaction are solved, and a natural, vivid and personalized interaction experience is achieved.
Patent Information
- Application Number
- CN202510638175.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-09-16
AI Technical Summary
In the existing technology, robots lack emotional expression when interacting with humans and the interaction form is mechanical, lacking naturalness and vividness.
By configuring the robot's target personality and using multimodal methods (voice, expression, action) for interaction, the personality can be updated in real time based on interactive feedback data to adapt to user needs and achieve dynamic adjustment.
It enhances the naturalness and vividness of robot interaction, meets personalized needs, realizes adaptive optimization, and improves the quality and effect of interaction.
Smart Images

Figure CN120645241A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to human-computer interaction technology, and in particular to a robot interaction processing method, device, equipment, storage medium and program product. Background Art
[0002] In human-robot interaction applications (i.e., human-computer interaction applications), relevant technologies use large language models and intelligent agent technologies to guide robots to interact with humans. However, this interaction is mostly text-based, and there are problems such as the lack of emotional expression of robots and the mechanical form of interaction. Summary of the Invention
[0003] The embodiments of the present application provide a robot interaction processing method, device, electronic device, computer-readable storage medium and computer program product, which can improve the naturalness and vividness of the robot interaction process.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] The present invention provides a method for processing robot interactions, including:
[0006] In response to a character configuration operation for a robot, configuring the character of the robot to a target character;
[0007] In response to an interactive operation directed to the robot, controlling the robot to interact in a multimodal manner that matches the target personality, the multimodal manner comprising at least two of the following: voice, expression, and action;
[0008] During the robot interaction process, the target personality is updated based on the interaction feedback data, and the robot is controlled to interact in a multimodal manner that matches the updated target personality.
[0009] The present invention provides a robot interaction processing device, including:
[0010] a configuration module, configured to configure the character of the robot to a target character in response to a character configuration operation on the robot;
[0011] an interaction module, configured to, in response to an interactive operation directed to the robot, control the robot to interact in a multimodal manner that matches the target personality, the multimodal manner comprising at least two of the following: voice, expression, and action;
[0012] An updating module is used to update the target personality based on interaction feedback data during the robot interaction process, and control the robot to interact in a multimodal manner that matches the updated target personality.
[0013] In the above scheme, the configuration module is also used to display personality editing prompt information in response to a trigger operation on the personality setting entrance associated with the robot; in response to the target content edited based on the personality editing prompt information, extract personality parameters from the target content, and configure the personality of the robot to the target personality corresponding to the personality parameters.
[0014] In the above scheme, before configuring the personality of the robot to the target personality corresponding to the personality parameters, the device also includes: a determination module, used to determine at least one preset personality associated with the robot; matching the personality parameters with each of the preset personalities to obtain a matching degree between the personality parameters and each of the preset personalities; and determining the preset personality whose matching degree exceeds a matching degree threshold as the target personality corresponding to the personality parameters.
[0015] In the above scheme, the configuration module is also used to determine the influence data that affects the personality generation in response to the trigger operation of the personality setting entrance associated with the robot by the target object, and the influence data includes at least one of the following: historical interaction data between the target object and the robot, the interaction scene between the target object and the robot, the object data of the target object, and the environmental information of the robot; based on the influence data, the personality of the robot is predicted to obtain the target personality, and the personality of the robot is configured to the target personality.
[0016] In the above scheme, the configuration module is also used to extract features from the influencing data to obtain key features corresponding to the influencing data, and the key features include at least one of the following: emotional features, scene features, and environmental features; based on the key features, personality parameters that affect personality generation are generated; personality mapping is performed based on the personality parameters to obtain the target personality corresponding to the personality parameters.
[0017] In the above scheme, the interaction module is also used to determine the second content of the robot's reply to the first content when the interaction operation indicates that the target object's interaction content with respect to the robot is the first content; perform semantic recognition and emotion recognition on the second content to obtain the target semantics and target emotion corresponding to the second content; in the mapping relationship under the target personality, determine the target expression associated with the target semantics based on the first mapping relationship between expression and semantics, determine the target action associated with the target semantics based on the second mapping relationship between expression and action, and determine the target timbre corresponding to the target emotion based on the third mapping relationship between emotion and timbre; control the robot to output the second content in the voice of the target timbre, and control the robot to synchronously execute the corresponding target expression and target action.
[0018] In the above scheme, the interactive operation is a shaking operation for the robot, and the interactive module is also used to control the robot to perform feedback animation for the shaking operation in a multimodal manner that matches the target personality when the robot is not in a standby state; when the robot is in a standby state and the shaking operation is not triggered on the robot within a target historical period, reset the shaking countdown for the robot, and when the shaking countdown returns to zero, control the robot to perform feedback animation for the shaking operation in a multimodal manner that matches the target personality.
[0019] In the above scheme, the interactive operation is a pressing operation of the function key of the robot. The interactive module is also used to control the robot to broadcast the introduction information of the robot in a multimodal manner that matches the target personality when the robot is in a normal state, and display the introduction information in the interactive interface of the robot; when the robot is in an abnormal state, control the robot to output prompt information for the abnormal state in a multimodal manner that matches the target personality, and the abnormal state includes at least one of the following: no signal, battery level is lower than the battery threshold, and membership expires.
[0020] In the above scheme, after the introduction information is displayed in the interactive interface of the robot, the device also includes: a binding prompt module, which is used to display the device serial number and graphic code of the robot in the interactive interface in response to a shaking operation on the robot; wherein, the graphic code is used for the identification terminal to download the robot's application when recognizing the graphic code, and bind the robot through the application.
[0021] In the above scheme, the update module is also used to collect interactive feedback data of the target object interacting with the robot, and the interactive feedback data includes at least one of the following: feedback expression, feedback timbre, feedback action, and the degree of recognition of the target personality of the robot; feature extraction is performed on the interactive feedback data to obtain feedback features of the target object; and personality parameters of the target personality are updated based on the feedback features.
[0022] In the above scheme, the update module is also used to determine the first scene in which the target object is located based on the feedback characteristics, and determine the second scene to which the target personality is adapted by default; when there is a conflict between the first scene and the second scene, the target personality is updated based on the first scene so that the updated target personality matches the first scene.
[0023] In the above scheme, after controlling the robot to interact in a multimodal manner that matches the updated target personality, the update module is also used to control the target personality of the robot to return from matching the first scene to matching the second scene when the target object exits the first scene.
[0024] An embodiment of the present application provides an electronic device, including:
[0025] a memory for storing computer-executable instructions or computer programs;
[0026] The processor is used to implement the robot interaction processing method provided in the embodiment of the present application when executing the computer executable instructions or computer program stored in the memory.
[0027] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions or a computer program for implementing the robot interaction processing method provided in the embodiment of the present application when executed by a processor.
[0028] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the robot interaction processing method provided in the embodiment of the present application is implemented.
[0029] The embodiments of the present application have the following beneficial effects:
[0030] By applying the embodiments of the present application, first, in response to the user's personality configuration operation, the robot can quickly switch to the target personality to meet the user's personalized needs in different scenarios; secondly, the robot interacts with the user in a multimodal manner that matches the target personality, including voice, expression, and action, which greatly enhances the naturalness and vividness of the interaction; finally, during the interaction process, the robot updates the target personality in real time based on the user's feedback data, and adjusts its multimodal interaction mode to achieve adaptive optimization. This dynamic adjustment mechanism based on feedback enables the robot to continuously learn and adapt to the user's preferences, thereby better meeting the user's needs and improving the quality and effect of the interaction. In short, through this dynamic configuration, multimodal interaction, and feedback-based update mechanism, the embodiments of the present application can significantly improve the robot's interactive performance, provide users with a more personalized, natural, and adaptable interactive experience, and thus promote the widespread application of robot technology in more fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 1 is a schematic diagram of the architecture of a robot interaction processing system 100 provided in an embodiment of the present application;
[0032] Figure 2is a structural diagram of an electronic device 500 provided in an embodiment of the present application;
[0033] Figure 3 1 is a flow chart of a robot interaction processing method provided in an embodiment of the present application;
[0034] Figure 4 2 is a schematic diagram of triggering a shaking operation provided in an embodiment of the present application;
[0035] Figure 5 This is a logical diagram of multimodal interaction provided by an embodiment of the present application;
[0036] Figure 6 This is a schematic diagram of the binding prompt interface provided in an embodiment of the present application. DETAILED DESCRIPTION
[0037] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0038] It is understandable that in the embodiments of the present application, when user information and other related data are involved, when the embodiments of the present application are applied to specific products or technologies, user permission or consent must be obtained, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards.
[0039] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0040] In the following description, the terms "first\second..." are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first\second..." can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0041] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0042] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0043] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0044] 1) Robots: In human-computer interaction applications, a robot is an intelligent system capable of perceiving its environment, making autonomous decisions, and executing tasks. It collects information through sensors (such as cameras, microphones, and sensor arrays), makes decisions using algorithms (such as machine learning and natural language processing), and interacts with humans through actuators (such as motors, displays, and speech synthesizers). Robots can be physical (such as service robots and industrial robots) or virtual (such as chatbots and virtual assistants).
[0045] Among them, service robots are robots used in homes, businesses or public places, and are designed to provide various services to humans, such as cleaning, moving, companionship, etc. Taking service robots as companion robots as an example, companion robots can perceive human emotions and voice commands through cameras and microphones, and provide companionship, entertainment and information query services. For example, it can have simple conversations with users, play music, and even provide navigation services to customers in shopping malls.
[0046] Chatbots are software-based virtual robots that use natural language processing technology to interact with users through text or voice, providing services such as information query and customer support. Virtual assistants are also software-based robots that use a graphical user interface or voice interaction to provide services such as calendar management, information query, and task reminders.
[0047] 2) Robot personality: This refers to a series of stable behavioral characteristics and emotional tendencies that a robot exhibits during its interactions with humans. These behavioral characteristics and emotional tendencies are expressed through the robot's language, behavior, expressions, and other means, enabling the robot to exhibit human-like personality traits when interacting with humans. Robot personality types include, but are not limited to, the following:
[0048] Lively and cheerful type: Robots with this personality show positive, optimistic, and enthusiastic behavioral characteristics, and are good at interacting with users in a humorous and relaxed manner. For example, a companion robot can interact with users in a shopping mall or home environment through cheerful voice tones, rich expressions, and lively movements. For example, the companion robot answers users' questions in a humorous way, and may even take the initiative to initiate some light-hearted topics to make users feel happy and relaxed. For another example, in a conversation, a chatbot can be configured to use a positive and optimistic language style and answer questions with a sense of humor. For example, when a user asks "What's the weather like today?", it may answer: "The weather is nice today, suitable for going out for a walk, but don't forget to bring a good mood!"
[0049] Rigorous and Professional: Robots with this personality demonstrate serious, rigorous, and professional behavior, focusing on detail and accuracy. They are suitable for formal settings or situations requiring specialized knowledge. For example, when providing services such as information search and calendar management, virtual assistants use a rigorous and professional language style. For example, when a user asks, "What are tomorrow's meetings?", the virtual assistant will clearly and accurately provide the specific meeting time and location. In another example, when answering user questions, the virtual assistant uses search engines to obtain accurate information and presents it to the user in a professional and objective manner. For example, when a user asks, "How do I treat a cold?", the virtual assistant will provide detailed medical advice and precautions. Furthermore, when handling important user tasks, the virtual assistant will display a serious and attentive attitude. For example, when a user needs to query important legal terms, the virtual assistant will provide relevant information in a serious and accurate manner.
[0050] Gentle and Friendly: Robots with this personality demonstrate gentle, friendly, and patient behavior, excelling at listening to and understanding users' needs and interacting with them in a gentle manner. For example, a virtual assistant might employ a gentle and friendly language style when interacting with users. For example, when a user is frustrated, it might say, "Don't worry, everything will be fine. How can I help you?" Another example is a companion robot, which displays a gentle and patient personality when accompanying the elderly or children. For example, a companion robot might tell stories in a soft voice and patiently answer children's questions, making them feel comfortable and at ease.
[0051] Humorous and playful: Robots with this personality type display humorous, playful, and witty behaviors, and are adept at entertaining users with humorous language and actions. For example, a chatbot might interact with users using humorous and playful language. For example, when a user asks, "What's your favorite color?" it might respond, "My favorite color is transparent because it makes me invisible." Another example is a companion robot, which uses humorous gestures and expressions to amuse users. For example, a companion robot might make funny expressions or gestures that make users laugh.
[0052] 3) In response, it is used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations executed can be real-time or have a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations executed are executed.
[0053] The embodiments of the present application provide a robot interaction processing method, device, electronic device, computer-readable storage medium and computer program product, which can improve the naturalness and vividness of the robot interaction process. The exemplary application of the electronic device provided by the embodiment of the present application is described below. The electronic device provided by the embodiment of the present application can be implemented as various types of user terminals such as laptop computers, tablet computers, desktop computers, set-top boxes, mobile devices (for example, mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable gaming devices), smart phones, smart speakers, smart watches, smart TVs, and vehicle-mounted terminals, and can also be implemented as servers. Below, an exemplary application when the device is implemented as a terminal will be described.
[0054] See also Figure 1 , Figure 1 This is a schematic diagram of the architecture of the robot interaction processing system 100 provided in an embodiment of the present application. In order to support an exemplary application, the terminal (terminal 400-1 and terminal 400-2 are shown as examples) is connected to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.
[0055] In some embodiments, the terminal can be the physical entity 400-1 of the robot (that is, the terminal itself is a robot, such as a service robot), and the terminal can also be a user-targeted 400-2 running an application that provides robot-related services (the robot in this case is a virtual robot, such as a chat robot, a virtual assistant, etc.). The server 200 is the background server corresponding to the robot, which can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal and the server can be directly or indirectly connected through wired or wireless communication, which is not limited in the embodiments of the present application.
[0056] In actual applications, the terminal responds to the character configuration operation for the robot by sending a character configuration request for the robot to the server 200; based on the character configuration request, the server 200 configures the robot's character to the target character, and responds to the interactive operation for the robot, controls the robot to interact in a multimodal manner that matches the target character, wherein the multimodal manner includes at least two of the following: voice, expression, and action; during the robot interaction process, the user's interactive feedback data is collected, and the target character is updated based on the interactive feedback data, and the robot is controlled to interact in a multimodal manner that matches the updated target character.
[0057] See also Figure 2 , Figure 2 This is a structural diagram of an electronic device 500 provided in an embodiment of the present application, with the electronic device 500 as an example. Figure 1 For example, Figure 2 The electronic device 500 shown includes: at least one processor 510, a memory 550, at least one network interface 520 and a user interface 530. The various components in the electronic device 500 are coupled together via a bus system 540. It is understood that the bus system 540 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 540 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 540 is not shown in FIG. Figure 2 Various buses are labeled as bus system 540 .
[0058] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0059] The memory 550 includes a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory. The memory 550 may optionally include one or more storage devices physically remote from the processor 510.
[0060] In some embodiments, the memory 550 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.
[0061] The operating system 551 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., which are used to implement various basic businesses and process hardware-based tasks; the network communication module 552 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 520. Exemplary network interfaces 520 include: Bluetooth, Wireless Compatibility Certification (WiFi), and Universal Serial Bus (USB), etc.
[0062] In some embodiments, the interactive processing device of the robot provided in the embodiments of the present application can be implemented in software. The interactive processing device of the robot provided in the embodiments of the present application can be provided as various software embodiments, including various forms including applications, software, software modules, scripts or codes. Figure 2 An interactive processing device 555 of a robot stored in a memory 550 is shown, which may be software in the form of programs and plug-ins, and includes a series of modules, including a configuration module 5551, an interactive module 5552, and an update module 5553. These modules are logical, and therefore may be arbitrarily combined or further split according to the functions implemented. The functions of each module will be described below.
[0063] In other embodiments, the device provided in the embodiments of the present application can be implemented in hardware. As an example, the device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the robot interaction processing method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs) or other electronic components.
[0064] In some embodiments, the terminal or server can implement the interactive processing method of the robot provided in the embodiment of the present application by running various computer executable instructions or computer programs. For example, computer executable instructions can be commands, machine instructions or software instructions at the microprogram level. The computer program can be a native program or software module in the operating system; it can be a local (Native) application (APPlication, APP), that is, a program that needs to be installed in the operating system to run, such as a robot control APP; it can also be a small program that can be embedded in any APP, that is, a program that can be run only by downloading it to a browser environment. In short, the above-mentioned computer executable instructions can be instructions in any form, and the above-mentioned computer program can be an application, module or plug-in in any form.
[0065] The following describes the robot interaction processing method provided by the embodiments of the present application with reference to the accompanying drawings. As previously mentioned, the electronic device 500 that implements the robot interaction processing method of the embodiments of the present application can be a terminal, a server, or a combination of the two. Therefore, the execution entity of each step will not be repeatedly described below.
[0066] The interactive processing method of the robot in the embodiment of the present application is described by taking the execution subject as a server as an example. Figure 3 , Figure 3 This is a flow chart of the robot interaction processing method provided by the embodiment of the present application, which will be combined with Figure 3 The steps shown are explained.
[0067] Step 101: In response to a character configuration operation for a robot, the character of the robot is configured as a target character.
[0068] In actual applications, before interacting with the robot, the user can pre-configure the robot's personality, such as configuring the robot's personality to the target personality. In this way, the robot can show the target personality characteristics of humans when interacting with the user.
[0069] In some embodiments, the robot's personality can be configured as a target personality in response to a personality configuration operation for the robot in the following manner: in response to a trigger operation on a personality setting entry associated with the robot, personality editing prompt information is output; in response to target content edited based on the personality editing prompt information, personality parameters are extracted from the target content, and the robot's personality is configured as the target personality corresponding to the personality parameters.
[0070] In practical applications, a clear personality setting entry can be displayed in the robot's user interface or interactive interface. This entry can be a button or a voice command trigger. The user can trigger the personality setting operation by clicking the button (e.g., clicking the "Personality Setting" button on the robot's control panel) or speaking a specific voice command (e.g., saying "Set personality" to the robot). After the user triggers the personality setting operation, a personality editing prompt can be provided to guide the user on how to edit the personality. This personality editing prompt can be in the form of text, voice, or a graphical interface, depending on the robot's interaction method. The personality description edited by the user based on the personality editing prompt can be a natural language instruction or a specific personality preference. For example, if the personality editing prompt is a text prompt, a personality editing prompt such as "Please enter the personality description you would like the robot to have" can be displayed in the robot's user interface or interactive interface. If the personality editing prompt is a voice prompt, the personality editing prompt can be a voice prompt spoken by the robot, such as "Please tell me what kind of personality you would like me to have." The user can then edit the personality description accordingly based on this personality editing prompt, such as "I need a serious English teacher" or "I tend to be extroverted."
[0071] After receiving the target content (i.e., personality description) input by the user, the server parses the target content and extracts key personality parameters. In actual implementation, the server can use a pre-trained language model to understand the user's intentions and extract specific personality parameters. After extracting the personality parameters corresponding to the target content, the robot's personality is configured to the target personality corresponding to the personality parameters based on the extracted personality parameters. This involves setting the robot's voice intonation, expression, action and other multimodal behaviors to match the target personality.
[0072] For example, when a user edits the target content as "Need a rigorous English teacher," the extracted personality parameters are "High rigor" and "Expertise: English teaching." The corresponding target personality characteristics are: smooth, clear voice, minimal body movements, and standard English pronunciation. Specifically, the robot speaks in a smooth, clear voice with standard pronunciation; expresses a serious expression with a focused gaze; and moves concisely, minimizing unnecessary body movements.
[0073] For example, if the user's target content is "Extroverted personality, likes to use metaphors," the extracted personality parameters "High extroversion" and "Prefers to use metaphors" are used. The corresponding target song for these personality parameters is characterized by cheerful and energetic voice, increased body movements, and the use of rich metaphors and vivid language. Specifically, in terms of voice, the robot speaks in a cheerful and energetic tone, using vivid language; in terms of facial expressions, the robot has rich expressions and bright eyes; and in terms of movement, the robot moves naturally and lively, with more gestures and body language.
[0074] Through the above method, users can define the robot's personality through natural language instructions or interaction examples. The server can extract personality parameters based on the user's input and configure the robot's personality to the target personality corresponding to the personality parameters. This flexible personality configuration method not only improves the user experience, but also enables the robot to better adapt to different user needs and interaction scenarios.
[0075] In some embodiments, before configuring the robot's personality to the target personality corresponding to the personality parameters, at least one preset personality associated with the robot may be determined first; the personality parameters and each preset personality are matched to obtain a matching degree between the personality parameters and each preset personality; and the preset personality whose matching degree exceeds a matching degree threshold is determined as the target personality corresponding to the personality parameters.
[0076] In actual applications, the server can also pre-define a set of personality templates (i.e., preset personalities) for the robot. Each personality template contains a set of specific personality parameters that describe the main characteristics of the corresponding preset personality. The preset personality can be designed based on common personality types (such as extroversion, introversion, rigor, humor, etc.) or specific scenarios (such as teachers, customer service, companion robots, etc.). For example, the preset personalities may include but are not limited to: extroversion (high extroversion, high humor, rich body movements), rigor (high rigor, steady voice and intonation, concise body movements), teacher (high rigor, high knowledge, clear voice and intonation), customer service (high patience, high friendliness, warm voice and intonation).
[0077] After extracting the corresponding personality parameters from the target content edited by the user, the server can match the extracted personality parameters with the preset personalities. This means matching the extracted personality parameters with the personality parameters associated with the preset personalities to determine the degree of match between the extracted personality parameters and each preset personality (i.e., the degree of match between the target content and the preset personality). The preset personality whose match exceeds a matching threshold (which can be set based on actual needs) is selected as the target personality that matches the target content. It should be noted that if the matching degrees of multiple preset personalities all exceed the matching threshold, the preset personality with the highest matching degree may be selected as the target personality.
[0078] For example, for the user input target content "Need a rigorous English teacher," the extracted personality parameters are: high rigor, professional field: English teaching. After matching these extracted personality parameters with the aforementioned multiple preset personalities, the target content matches the preset personality of "teacher" at a degree of 0.8, and matches the preset personality of "extrovert" at a degree of 0.2. Assuming a matching threshold of 0.6, the preset personality of "teacher" is determined as the target personality matching the target content. After determining the target personality, the robot's personality can be configured to the target personality (i.e., the preset personality of "teacher" described above). Based on the target personality, a corresponding personality template (i.e., the personality template of "teacher") is loaded. The personality template includes specific multimodal behavioral parameters such as speech, facial expression, and movement. Based on the personality template corresponding to the target personality, the robot's voice intonation, facial expression, and movement are adjusted to embody the target personality, for example, adjusting the robot's voice intonation to be clear and steady, its facial expression to be serious and focused, and its movements to be concise and standardized.
[0079] For example, for the user input target content "personality tends to be extroverted, likes to use metaphors," personality parameters are extracted: high extroversion, preference for metaphors. After matching these extracted personality parameters with the aforementioned preset personalities, the target content matches the preset personality of "extrovert" at a degree of 0.9, the preset personality of "humorous" at a degree of 0.6, and the preset personality of "teacher" at a degree of 0.3. Assuming the matching threshold of 0.6, the preset personality of "extrovert" (with the highest matching degree and exceeding the threshold) is determined as the target personality matching the target content. After determining the target personality, the robot's personality can be configured to the target personality (i.e., the preset personality of "extrovert"), and the corresponding personality template (i.e., the personality template of "extrovert") is loaded based on the target personality. The robot's voice and tone are adjusted to be cheerful and energetic, with rich and vivid expressions, natural and lively movements, and frequent use of metaphors in its language.
[0080] Through the above method, users can define the robot's personality through natural language instructions. The system can extract personality parameters based on the user's input and determine the target personality by matching the preset personality template. This method not only improves the flexibility and accuracy of personality configuration, but also supports "personality cloning", enabling the robot to better adapt to different user needs and interaction scenarios.
[0081] In some embodiments, the robot's personality can be configured as a target personality in response to a personality configuration operation for the robot in the following manner: in response to a trigger operation of a personality setting entry associated with the robot by a target object, influence data that affects the generation of the personality is determined, wherein the influence data includes at least one of the following: historical interaction data between the target object and the robot, interaction scenes between the target object and the robot, object data of the target object, and environmental information of the robot; the robot's personality is predicted based on the influence data to obtain a target personality, and the robot's personality is configured as the target personality.
[0082] Here, influencing data refers to various information that can influence the generation of a robot's personality, including historical interaction data, interaction scenarios, target object data, and the robot's environmental information. Specifically, all historical interactions between the user (i.e., target object) and the robot are recorded, including conversation content (conversation content can be recorded in text form), feedback scores (feedback scores can be user satisfaction scores for the robot's service), and interaction frequency (interaction frequency can be the number of times the user interacts with the robot per unit time). Interaction scenarios describe the specific contexts in which the user interacts with the robot, such as home, office, or school. Alternatively, interaction scenarios describe the user's circumstances, such as a fall or successful challenge. Target object data refers to basic user information, such as age, gender, occupation, upbringing, intelligence, education, interests, and hobbies. Environmental information refers to the environment in which the robot or target object is located, such as time, location, and weather (available through the weather interface). Alternatively, it refers to external information that may influence the prediction of the robot's personality, such as news events (available through the news aggregation interface).
[0083] In some embodiments, the robot's personality can be predicted based on the influence data to obtain a target personality in the following manner: feature extraction is performed on the influence data to obtain key features corresponding to the influence data, wherein the key features include at least one of the following: emotional features, scene features, and environmental features; based on the key features, personality parameters that influence personality generation are generated; personality mapping is performed based on the personality parameters to obtain a target personality corresponding to the personality parameters.
[0084] In practical applications, after collecting the aforementioned influence data, it can be preprocessed, such as through cleaning and normalization, before personality prediction is based on the preprocessed influence data. To predict personality from this preprocessed influence data, key features can be extracted from the influence data. These key features reflect the primary characteristics and patterns of the influence data. For robot personality prediction, key features include emotional features, scene features, and environmental features. Emotional features reflect the emotional tendencies of users interacting with the robot, such as positive, negative, or neutral. Natural language processing techniques, such as sentiment analysis models (e.g., BERT), can be used to perform sentiment analysis on conversations in historical interaction data to extract emotional features. For example, if a user's conversation with the robot is "The weather is great today, and I'm in a good mood," the sentiment analysis model outputs a sentiment tendency of "positive" with a confidence level of 0.9. Scene features describe the specific context in which the user interacts with the robot, such as home, office, or school. When extracting scene features from interaction scene data, predefined scene labels can be used or inferred from contextual information (e.g., location and time). For example, if the user interacts with the robot in a home environment, the extracted scene feature is "home." Environmental features describe the robot's surroundings, such as time, location, and weather. Specific time, location, and weather features are extracted from the environmental information data, such as the current time is 8 p.m., the location is the living room, and the weather is sunny.
[0085] After extracting key features, corresponding personality parameters can be generated based on these key features. Personality parameters are specific numerical values or labels that describe the robot's personality, such as extroversion, conscientiousness, and humor. When mapping emotional features to personality parameters, the personality parameters can be adjusted based on the type and confidence level of the emotional feature. For example, if the emotional tendency or type indicated by the emotional feature is "positive" and the confidence level is 0.9, it can be mapped to personality parameters of high humor (humor 0.8) and high extroversion (extroversion 0.7). When mapping scene features to personality parameters, the personality parameters can be adjusted based on the scene type. For example, a family scene can be mapped to personality parameters of high extroversion (extroversion 0.8) and high humor (humor 0.7), while an office scene can be mapped to personality parameters of high conscientiousness (conscientiousness 0.9) and low humor (humor 0.3). When mapping environmental characteristics to personality parameters, the personality parameters may be adjusted according to specific values of the environmental characteristics. For example, 8 o'clock in the evening may be mapped to high humor (humor 0.7), and a sunny day may be mapped to high extroversion (extroversion 0.8).
[0086] Personality mapping based on personality parameters involves mapping the generated personality parameters to a specific target personality, enabling the robot to exhibit the corresponding personality traits. During the mapping process, personality parameters extracted from emotional, scene, and environmental characteristics are first integrated, such as using weighted averaging or other fusion methods to combine personality parameters from different sources to obtain comprehensive personality parameters. For example, if the personality parameters from the emotional feature mapping are humor 0.8 and extroversion 0.7, the personality parameters from the scene feature mapping are extroversion 0.8 and humor 0.7, and the personality parameters from the environmental feature mapping are humor 0.7 and extroversion 0.8, then the resulting comprehensive personality parameters are humor 0.75 and extroversion 0.75. When generating a target personality based on the comprehensive personality parameters, the comprehensive personality parameters are matched against a preset personality template, and the closest matching personality template is selected as the target personality. For example, if the comprehensive personality parameters are humor 0.75, extroversion 0.75, and conscientiousness 0.3, the "extroversion" template that best matches the preset personality templates is selected as the target personality.
[0087] In practical applications, multidimensional personality vectors can be predefined, such as extraversion, conscientiousness, and humor. Each dimension can be represented by a numerical value, and the numerical range can be set according to specific needs, such as a floating-point number between 0 and 1. For example, extraversion can be represented as a value between 0 (very introverted) and 1 (very extraverted); conscientiousness can be represented as a value between 0 (very casual) and 1 (very conscientious); and humor can be represented as a value between 0 (very serious) and 1 (very humorous). Personality parameters are then initialized based on key features extracted from the influencing data. For example, heuristic rules or simple statistical methods can be used for initialization. Extraversion and humor can be preliminarily determined based on the emotional orientation of the conversation content, and extraversion can be preliminarily determined based on the frequency of interaction. For example, if the conversation contains more positive emotional words, humor is preliminarily determined to be higher; if the frequency of interaction is higher, extraversion is preliminarily determined to be higher. Then, based on the sentiment analysis results and other input data (such as feedback scores, interaction frequency, and environmental context), personality parameters are updated. For example, if the sentiment tendency is "positive" and the confidence level is high, the humor and extraversion parameters can be appropriately increased. If the feedback score is high and the interaction frequency is high, the extraversion parameter can be further increased. If the scene type is a social scene, the humor parameter can be appropriately increased. Finally, methods such as weighted averaging can be used to comprehensively consider the impact of different data on personality parameters. For example, the weight of the sentiment analysis result can be set to 0.6, the weight of the feedback score to 0.3, and the weight of the interaction frequency to 0.1, and then the personality parameters can be updated based on these weights.
[0088] In practical applications, target personalities are also defined based on business needs and user preferences. For example, in office settings, the target personality might favor rigor and professionalism; in social settings, it might favor extroversion and humor. The target personality can be represented as a multidimensional vector, with the value of each dimension representing the target personality's expected value in that dimension. For example, the target personality vector could be represented as [0.7, 0.9, 0.5], representing an extroversion of 0.7, a conscientiousness of 0.9, and a humor of 0.5, respectively. When mapping the updated personality parameters to the target personality, methods such as interpolation and smoothing can be used. For example, if the personality parameters differ significantly from the target personality in a certain dimension, the value of that dimension can be gradually adjusted to bring it closer to the target personality. For example, if the updated personality parameters are [0.6, 0.8, 0.4], they can be mapped to the target personality [0.7, 0.9, 0.5].
[0089] After determining the target personality, a corresponding personality description or behavior pattern can be generated. For example, if the target personality is extroverted, conscientious, and humorous, a personality description can be generated: "This robot is cheerful, conscientious, and enjoys communicating with people through humor." The target personality can be output as a vector or a structured description, depending on the application scenario and requirements. After determining the target personality, a target personality template is loaded based on the target personality to adjust the robot's multimodal behaviors, such as voice intonation, facial expressions, and movements. For example, if the target personality is "extroverted," an extrovert personality template is loaded to adjust the robot's voice intonation to be cheerful and energetic, its facial expressions to be rich and lively, and its movements to be natural and lively.
[0090] Through the above method, users can trigger the personality setting entrance, and the system will predict the target personality based on influencing data such as historical interaction data, interaction scenarios, target object data and environmental information, and configure the robot's personality to the target personality. This method not only improves the flexibility and accuracy of personality configuration, but also can dynamically adjust the robot's personality according to different user needs and interaction scenarios, thereby improving user experience.
[0091] Step 102: In response to an interactive operation directed to the robot, control the robot to interact in a multimodal manner that matches the target personality, wherein the multimodal manner includes at least two of the following: voice, expression, and action.
[0092] In some embodiments, the robot can be controlled to interact in a multimodal manner that matches the target personality in the following manner: when the interactive operation indicates that the target object's interactive content with respect to the robot is the first content, the second content of the robot's reply to the first content is determined; semantic recognition and emotion recognition are performed on the second content to obtain the target semantics and target emotion corresponding to the second content; in the mapping relationship under the target personality, the target expression associated with the target semantics is determined based on the first mapping relationship between expression and semantics, the target action associated with the target semantics is determined based on the second mapping relationship between expression and action, and the target timbre corresponding to the target emotion is determined based on the third mapping relationship between emotion and timbre; the robot is controlled to output the second content in a voice with the target timbre, and the robot is controlled to synchronously execute the corresponding target expression and target action.
[0093] In actual applications, user interactions with the robot can be voice commands, text input, or other forms of interaction. The server generates a second response based on the user's first response to the robot. For example, if the user first asks the robot, "Robot, what's the weather like today?", the server determines the robot's second response to be, "It's sunny today, perfect for going out." After determining the second response, the server performs semantic and sentiment recognition on the response. For example, natural language processing techniques are used to perform semantic analysis on the second response to extract key semantic information, and sentiment analysis models are used to perform sentiment analysis on the second response to understand its meaning and emotional orientation. For example, semantic recognition of the second response, "It's sunny today, perfect for going out," extracts the target semantics "sunny" and "perfect for going out," while sentiment recognition yields the target sentiment "positive."
[0094] Mapping relationships map the semantics, emotions, and other information associated with the target personality into specific multimodal expressions (such as speech, facial expressions, and actions). These mapping relationships include a first mapping relationship between facial expressions and semantics, a second mapping relationship between facial expressions and actions, and a third mapping relationship between emotions and timbre. Continuing with the example above, let's assume the target personality is "extrovert," characterized by high extroversion and a high sense of humor. The target semantics are "sunny weather" and "suitable for going out," and the target emotion is "positive." When determining multimodal expressions based on the first mapping relationship, the target facial expressions are determined based on the first mapping relationship between facial expressions and semantics. For example, for target semantics like "sunny weather" and "suitable for going out," an extrovert might display the target facial expressions of "smiling" and "excited." The target actions are determined based on the second mapping relationship between facial expressions and actions. For example, an extrovert might display the gestures of "waving" and "nodding." The target timbre is determined based on the third mapping relationship between emotion and timbre. For example, a positive emotion and an extrovert might select a "cheerful" timbre. In this way, it depends on controlling the robot to interact in a multimodal manner. The robot is controlled to interact with the user through various means such as voice, expression and action. Specifically, the robot is controlled to output the reply content in the target tone. For example, the robot says in a cheerful tone: "The weather is sunny today, suitable for going out". At the same time, the robot's expression system is controlled to make it show the target expression. For example, when the robot says in a cheerful tone: "The weather is sunny today, suitable for going out", it smiles and shows an excited expression; at the same time, the robot's action system is also controlled to make it perform the target action. For example, based on the above, the robot waves and nods to the user.
[0095] Through the above method, the robot can interact in a multimodal manner (voice, expression, action) according to the user's interactive operations and target personality. This method not only makes the robot's interaction more natural and vivid, but also can flexibly adjust the interaction method according to different personality characteristics and scenario requirements, thereby improving the user experience.
[0096] In some embodiments, when the interactive operation is a shaking operation for the robot, the robot can be controlled to interact in a multimodal manner that matches the target personality in the following manner: when the robot is not in a standby state, the robot is controlled to perform feedback animation for the shaking operation in a multimodal manner that matches the target personality; when the robot is in a standby state and the shaking operation is not triggered for the robot within the target historical period, the shaking countdown of the robot is reset, and when the shaking countdown returns to zero, the robot is controlled to perform feedback animation for the shaking operation in a multimodal manner that matches the target personality.
[0097] In practice, a shake action refers to a user shaking a robot, typically detected by an accelerometer or gyroscope. When a user triggers a shake action on the robot, the accelerometer or gyroscope collects motion data. When the acceleration exceeds a certain value or the angle changes beyond a preset threshold for a shake action, a shake action is detected, triggering the corresponding interaction logic.
[0098] See also Figure 4 , Figure 4 This is a schematic diagram of triggering the shaking operation provided in the embodiment of the present application, such as Figure 4 As shown, in step 201, the state of the robot is determined.
[0099] In step 202 , a shaking operation for the robot is received.
[0100] In step 203, it is determined whether the robot is in a standby state. A non-standby state indicates that the robot is "idle", in which case step 206 is executed. A standby state indicates that the robot is executing other interactive logic, in which case step 204 is executed.
[0101] In step 204 , it is determined whether a shaking operation has been triggered within the target historical period.
[0102] Here, when the robot is in standby mode (i.e., in interactive mode), it means that the robot is executing other interactive logic. In order to avoid the shaking operation interrupting the interactive logic being executed by the robot, the robot is not immediately controlled to execute the feedback animation. Instead, it detects whether the shaking operation has been triggered within the target historical period (which can be set according to actual needs, such as set to 1 hour). If the robot is in standby mode and the shaking operation has not been triggered within the target historical period, step 205 is executed.
[0103] In step 205, the shaking countdown is reset. The shaking countdown can be set according to actual needs, such as 10 minutes or 1 hour.
[0104] In step 206 , the robot is controlled to perform feedback animation.
[0105] Here, when the robot is in an off-standby state (i.e., in non-interactive mode), or when the shaking countdown returns to zero, the robot is controlled to perform feedback animation, that is, the robot is controlled to perform feedback animation for the shaking operation in a multimodal manner according to the target personality. For example, if the target personality is "extrovert", the robot may say in a cheerful voice: "Hey, you shook me!" while showing an excited expression and waving gesture.
[0106] For example, if the target personality is "humorous," the robot might randomly play one or more preset animations, such as a funny animation and a humorous voice saying, "Shake me again! I'm going to make you laugh!" Another example is when the robot's emotional tendency is negative, a series of feedback animations might be played, such as the robot expressing sadness, shock, excitement, anger, violent shaking, collapse (crying), vomiting, fatigue, and sticking out its tongue from exhaustion.
[0107] Through the above method, the robot can respond to shaking operations in a multimodal manner according to different states and target personalities. This method not only makes the robot's interaction more natural and vivid, but also can flexibly adjust the interaction method according to different personality characteristics and scene requirements, thereby improving the user experience.
[0108] In some embodiments, when the interactive operation is a pressing operation on a function key of the robot, the robot can be controlled to interact in a multimodal manner that matches the target personality in the following manner: when the robot is in a normal state, the robot is controlled to broadcast introduction information about the robot in a multimodal manner that matches the target personality, and the introduction information is displayed in the robot's interactive interface; when the robot is in an abnormal state, the robot is controlled to output prompt information for the abnormal state in a multimodal manner that matches the target personality, and the abnormal state includes at least one of the following: no signal, battery level is lower than a battery threshold, and membership expired.
[0109] In actual applications, function keys can be volume control keys (or volume plus or minus keys), menu keys, etc. When the user presses a specific function key on the robot, it can usually trigger a specific function or information display. For example, when a pressing operation is received on the robot, it detects whether the robot's current state is normal or abnormal. When the robot's current state is normal (i.e., there is a signal, the battery level is higher than the battery level group, the membership has not expired, etc.), the robot is controlled to interact with the user through various methods such as voice, expression and action. For example, according to the target personality, the robot is controlled to broadcast the robot's introduction information in a multimodal manner. Taking the target personality as friendly as an example, the robot says in a gentle voice: "Hi, I am your smart assistant. I can help you complete various tasks, such as querying information, setting reminders, etc.", and displays text in the robot's interactive interface: "I am your smart assistant. I can help you complete various tasks, such as querying information, setting reminders, etc." When the robot's current state is abnormal (i.e., no signal, battery level below a certain level, membership expiration, etc.), the control robot outputs a multimodal prompt message tailored to the target personality. For example, if the target personality is humorous and the abnormal state is caused by low battery, the control robot will say in a humorous voice, "Oops, my battery is almost out! I need to charge quickly or I'll go on strike!" and display the text "Battery low, please charge now" on the robot's interface. For another example, if the target personality is professional and the abnormal state is caused by membership expiration, the control robot will say in a professional voice, "Your membership has expired. To continue enjoying our advanced features, please renew your membership promptly." and display the text "Your membership has expired. Please renew your membership promptly." on the robot's interface.
[0110] Take the volume up and down keys as an example, see Figure 5 , Figure 5 This is a logical diagram of a multimodal interaction provided by an embodiment of the present application. The method includes:
[0111] In step 301, the state of the robot is determined.
[0112] In step 302 , a pressing operation on the robot is received.
[0113] In step 303, it is determined whether the robot is in an abnormal state.
[0114] Here, when the robot is in a normal state, step 304 is executed; when the robot is in an abnormal state, any one of steps 305A to 205C is executed.
[0115] In step 304, introductory information is displayed.
[0116] In actual applications, when the robot is in a normal state, it can automatically initiate a dialogue request such as "Please introduce yourself". Then, when the robot receives the self-introduction content returned by the member's cloud-based language model, it automatically makes the corresponding voice broadcast.
[0117] In step 305A, a signal abnormality prompt is displayed.
[0118] Here, when the robot's abnormality is caused by a signal abnormality (such as no signal), in addition to displaying the robot's self-introduction information in the robot's interactive interface, a voice broadcast prompt such as "no signal, no signal" is also output.
[0119] In step 305B, a power abnormality prompt is displayed.
[0120] Here, when the robot's abnormality is caused by low battery, in addition to displaying the robot's self-introduction information in the robot's interactive interface, a voice broadcast prompt such as "Please charge me" is also output.
[0121] In step 305C, a membership expiration reminder is displayed.
[0122] Here, when the robot's abnormality is caused by membership expiration, in addition to displaying the robot's self-introduction information in the robot's interactive interface, a voice broadcast prompt such as "The cute baby cannot pass through the human world" is also output.
[0123] Through the above method, the robot can respond to function key pressing operations in a multimodal manner according to different states (normal state or abnormal state) and target personality. This method not only makes the robot's interaction more natural and vivid, but also can flexibly adjust the interaction method according to different personality characteristics and scene requirements, thereby improving the user experience.
[0124] In some embodiments, after the introduction information is displayed in the robot's interactive interface, the robot's device serial number and graphic code are displayed in the interactive interface in response to a shaking operation on the robot; wherein, the graphic code is used for the identification terminal to download the robot's application and bind the robot through the application when recognizing the graphic code.
[0125] See also Figure 6 , Figure 6This is a schematic diagram of the binding prompt interface provided by an embodiment of the present application. After the robot's self-introduction information is displayed in the interactive interface, the user shakes the robot. Upon receiving the shaking operation directed to the robot, the robot's device serial number and graphic code are displayed in the interactive interface. The device serial number is the robot's unique identifier, used to identify and bind the robot. The graphic code is typically a QR code or barcode for recognition by an identification terminal (such as a mobile phone). When the identification terminal (such as a mobile phone) recognizes the graphic code, it automatically jumps to the app store to download the robot's application. The user uses the downloaded application to bind the robot using the device serial number to control the robot through the identification terminal.
[0126] Through the above method, the robot can display introductory information after the user presses the function key, and display the device serial number and graphic code after the user shakes the robot. The graphic code is used for identification by the identification terminal, downloading the robot's application, and binding the robot through the application. This method not only improves the user experience, but also simplifies the robot binding and management process, allowing users to use the robot's functions more conveniently.
[0127] Step 103: During the robot interaction process, the target personality is updated based on the interaction feedback data, and the robot is controlled to interact in a multimodal manner that matches the updated target personality.
[0128] In some embodiments, the target personality can be updated based on the interactive feedback data in the following manner: collecting interactive feedback data of the target object interacting with the robot, wherein the interactive feedback data includes at least one of the following: feedback expression, feedback timbre, feedback action, and the degree of recognition of the target personality of the robot; performing feature extraction on the interactive feedback data to obtain feedback features of the target object; and updating the personality parameters of the target personality based on the feedback features.
[0129] In practical applications, during the interaction between a robot and a target object, the robot's multimodal interaction methods can be adjusted and updated based on the target object's interactive feedback data. Interaction feedback data refers to the various feedback information expressed by the target object (user) during the interaction with the robot. This information can reflect the user's acceptance of and preference for the robot's personality. When collecting interaction feedback data, the robot's camera captures the user's facial expressions. Expression recognition technology (such as a deep learning-based facial expression recognition model) identifies the user's facial expressions, such as smiles and frowns. The robot's microphone captures the user's voice. Speech analysis technology (such as a deep learning-based voice sentiment analysis model) analyzes the user's timbre and determines their emotional tendencies, such as happiness and anger. The robot's sensors (such as accelerometers and gyroscopes) capture the user's movements. Motion recognition technology (such as a deep learning-based motion recognition model) identifies the user's movements, such as nodding and shaking their heads. When collecting the user's approval of the target personality, questionnaires, direct inquiries, indirect observations, or user feedback can be used to collect user approval of the robot's target personality. For example, users can directly give ratings or feedback, or if it is detected that users have given positive feedback on the robot's humorous expressions many times, the humor parameter will be gradually increased.
[0130] Feature extraction is the process of extracting useful information from interactive feedback data. This information can reflect the user's feedback characteristics on the robot's personality. Emotional features, such as happiness, sadness, and anger, are extracted from feedback expressions. For example, if a user smiles, the emotional feature extracted is "happy." Emotional features, such as happiness, anger, and calmness, are extracted from the feedback timbre. For example, if the user's voice timbre is high and the speaking speed is fast, the emotional feature extracted is "happy." Action features, such as nodding, shaking head, and waving, are extracted from feedback actions. For example, if a user nods, the action feature extracted is "affirmative." Approval features, such as high, medium, and low, are extracted from the user's degree of approval of the robot's target personality. For example, if a user gives a high rating, the approval feature extracted is "high."
[0131] When updating the target personality's personality parameters based on feedback features, the parameters are adjusted accordingly, allowing the robot to better adapt to the user's preferences and feedback. For example, when updating emotion parameters, if the user's emotional feedback is "happy" or "delighted," the target personality's "humor" and "extroversion" parameters are increased. If the user's emotional feedback is "angry" or "sad," the "humor" and "extroversion" parameters are decreased, and the "conscientiousness" parameter is increased. When updating action parameters, if the user's action feedback is "nodding" or "waving," the "extroversion" and "liveliness" parameters are increased. If the user's action feedback is "shaking head" or "frowning," the "extroversion" and "liveliness" parameters are decreased, and the "conscientiousness" parameter is increased. When updating approval parameters, if the user's approval of the target personality is "high," the current parameters of the target personality remain unchanged. If the user's approval of the target personality is "low," the parameters of the target personality are adjusted based on the user's feedback to better align with the user's preferences.
[0132] For example, assuming that the user is satisfied with the robot's humorous performance, and the user smiles and nods during the interaction with the robot, the emotional feature extracted from the feedback expression is "happy", and the action feature extracted from the feedback action is "affirmative". When the personality parameters of the target personality are updated based on the feedback features, the "humor" and "extraversion" parameters of the target personality are added to make the updated target personality more humorous and extroverted. In this way, when controlling the robot to interact with the updated target personality, the robot says in a cheerful voice: "I am glad that you like my humor so much. I will continue to work hard to make you happy!" At the same time, the robot shows a smiling expression and a waving action.
[0133] For another example, suppose the user expresses dissatisfaction with the robot's serious performance. During the interaction with the robot, the user frowns and shakes his head. The emotional feature extracted from the feedback expression is "anger", and the action feature extracted from the feedback action is "negation". When updating the personality parameters of the target personality based on the feedback features, the "rigorousness" parameter of the target personality is reduced, and the "humor" and "extraversion" parameters are increased, making the updated target personality more humorous and extroverted. In this way, when the robot is controlled to interact with the updated target personality, it says in a gentle voice: "It seems that I was too serious just now. Let me change my way and relax." At the same time, the robot shows a smiling expression and a waving action.
[0134] Through the above method, the robot can dynamically update the target personality according to the user's feedback data during the interaction process, and interact in a multimodal manner that matches the updated target personality. This method not only makes the robot's interaction more natural and vivid, but also can flexibly adjust the interaction method according to the user's real-time feedback, thereby improving the user experience.
[0135] In some embodiments, the target personality can be updated based on the feedback features in the following manner: determining the first scene in which the target object is located based on the feedback features, and determining the second scene to which the target personality is adapted by default; in the event of a conflict between the first scene and the second scene, updating the target personality based on the first scene so that the updated target personality matches the first scene.
[0136] In practical applications, the first scenario refers to the actual scenario the target subject (user) is currently in. This scenario is determined based on feedback features, such as those extracted from feedback expression, feedback timbre, and feedback action data. A pre-trained scenario recognition model (e.g., a deep learning-based scenario recognition model) analyzes these feedback features to determine the user's first scenario. For example, facial expression and action recognition can be used to determine whether the user is in an emergency (e.g., falling). The second scenario refers to the scenario for which the target personality is default-adapted, typically preset during personality configuration. For example, if the target personality is "humorous," the default scenario is "leisure and entertainment." By matching the first and second scenarios, it is determined whether there is a conflict. A scenario conflict occurs when the first and second scenarios do not match, requiring the target personality to be adjusted to accommodate the first scenario. Specifically, if there is a conflict between the first and second scenarios, the target personality is updated based on the first scenario. For example, if the first scenario is "urgent task" and the second scenario is "leisure and entertainment," the target personality's personality parameters are adjusted based on the characteristics of the first scenario. For example, in the emergency task scenario, the "humorousness" is reduced while the "rigor" and "responsiveness" are increased. In conflicting scenarios, the scene status determines whether the default personality behavior is overridden. When the sensor detects that the user is in an emergency state (such as falling), the emergency response mechanism is triggered. For example, even if the default personality is "calm", rapid acceleration and high-pitched reminders are still triggered. That is, by setting priority rules, emergency scenarios take precedence over default personality behaviors. For example, emergency mission scenarios have higher priority than leisure and entertainment scenarios.
[0137] For example, a user interacting with a robot in an emergency task scenario exhibits a frown and rapid speech. Through facial expression recognition and speech analysis, the user's first scenario is determined to be "Emergency Task." Assuming the target personality is "Humorous," the default second scenario is "Leisure and Entertainment." Comparing the first and second scenarios reveals a conflict. Based on the characteristics of the first scenario, the "Humor" level is reduced, while the "Strictness" and "Response Speed" are increased. Because the Emergency Task scenario takes precedence over the Leisure and Entertainment scenario, the robot's personality is adjusted to "Strictness." Consequently, when interacting with the robot using the updated target personality, the robot can be controlled to speak in a smooth, clear voice, saying, "Please tell me the specific problem. I'll help you solve it as soon as possible." The robot also displays a focused expression and quick movements.
[0138] For example, consider a user interacting with a robot in an emergency situation. Sensors detect the user falling, with a distressed expression and rapid speech. Using sensor data and facial recognition, the robot determines the user's first scenario as "Emergency." Assuming the target personality is "Calm," the default second scenario is "Daily Communication." Comparing the first and second scenarios reveals a conflict. Based on the characteristics of the first scenario, "Response Speed" and "Emergency Alert" are added. Because the emergency scenario takes priority over the daily communication scenario, the robot triggers the emergency response mechanism. Therefore, when controlling the robot to interact with the updated target personality, it rapidly accelerates to the user's side, saying in a high-pitched voice, "Did you fall? I need urgent help. Please let me know!" The robot also displays concern and moves quickly.
[0139] Through the above method, the robot can dynamically update the target personality according to the user's feedback characteristics during the interaction process, and adjust its behavior through a dynamic priority mechanism in conflict scenarios. This method not only makes the robot's interaction more natural and vivid, but also can flexibly adjust the interaction method according to the user's real-time feedback and scenario requirements, thereby improving the user experience.
[0140] In some embodiments, after the robot is controlled to interact in a multimodal manner that matches the updated target personality, when the target object exits the first scene, the target personality of the robot is restored from matching the first scene to matching the second scene.
[0141] In actual applications, the target object's exit from the first scene refers to the user leaving the current actual scene (such as an emergency task, leisure and entertainment, etc.) and returning to the default scene. During the interaction process, the target object's feedback characteristics (such as expression, timbre, movement, etc.) and environmental information (such as sensor data) are continuously detected. When the feedback characteristics and environmental information indicate that the user is no longer in the first scene, it is determined that the user has exited the first scene. For example, when the user leaves the emergency task scene, the expression and voice return to normal, and the sensor data no longer indicates an emergency state.
[0142] When the target object exits the first scene, the robot's target personality is controlled to be restored to match the second scene. That is, after the target object exits the first scene, the robot's target personality is restored to the default personality that matches the second scene. For example, when restoring the personality parameters, the personality parameters of the target personality are adjusted according to the characteristics of the second scene to restore it to the default state. For example, if the second scene is "leisure and entertainment", the target personality is restored to "humorous". In this way, the robot can be controlled to perform multimodal interaction with the restored target personality, including voice, expression, and movement. For example, after returning to "humorous", the robot says in a cheerful voice: "Mission completed, now you can relax!"
[0143] For example, a user exits an urgent task scenario. During their interaction with the robot, the user displays a frown and speaks in a hurried voice. Through facial expression recognition and voice analysis, the user's first scenario is determined to be "urgent task." Assuming the target personality is "humorous," the default second scenario is "leisure and entertainment." Comparing the first and second scenarios reveals a conflict. Based on the characteristics of the first scenario, the "humorousness" is reduced, while the "rigorousness" and "responsiveness" are increased. Because the urgent task scenario takes priority over the leisure and entertainment scenario, the robot's personality is adjusted to "rigorous." Therefore, when controlling the robot to interact with the updated target personality, the robot can say in a smooth, clear voice, "Please tell me the specific problem. I'll help you solve it as soon as possible." Simultaneously, the robot displays a focused expression and quick movements. When the user completes the urgent task, their facial expression and voice return to normal, and the sensor data no longer indicates an urgent state. The target is then considered to have exited the first scenario. The target personality is then restored to match the second scenario. For example, if the target personality is restored to "humorous," the robot can say in a cheerful voice, "Mission completed. Now you can relax!" Simultaneously, the robot displays a smile and relaxed movements.
[0144] For example, consider a user exiting an emergency scenario. Sensors detect a fall, with a distressed expression and rapid speech. Based on sensor data and facial recognition, the user's first scenario is determined to be an "emergency scenario." Assuming the target personality is "calm," the default second scenario is "daily communication." Comparing the first and second scenarios reveals a conflict. Based on the characteristics of the first scenario, "response speed" and "emergency reminder" are added. Because the emergency scenario takes precedence over the daily communication scenario, the robot triggers the emergency response mechanism. Therefore, when the robot is controlled to interact with the updated target personality, it rapidly accelerates to the user's side and says in a high-pitched voice, "Did you fall? I need urgent help. Please let me know!" The robot displays a concerned expression and quick movements. When the user recovers from the emergency, their expression and speech return to normal, and the sensor data no longer indicates an emergency state, the target is considered to have exited the first scenario. The target personality is then restored to match the second scenario. For example, if the target personality is restored to "calm," the robot is controlled to say in a smooth, clear voice, "Are you feeling better now? Is there anything else I can do?" with a gentle expression and steady movements.
[0145] Through the above method, the robot can automatically restore the target personality to the default personality matching the second scene after the target object exits the first scene. This method not only makes the robot's interaction more natural and vivid, but also can flexibly adjust the interaction method according to the user's real-time scene requirements, thereby improving the user experience.
[0146] So far, the robot interaction processing method provided by the embodiment of the present application has been described in conjunction with the exemplary application and implementation of the electronic device provided by the embodiment of the present application. The following will continue to describe the various modules in the robot interaction processing device 555 provided by the embodiment of the present application that cooperate to implement the robot interaction processing scheme. Configuration module 5551 is used to configure the robot's personality to a target personality in response to a personality configuration operation for the robot; interaction module 5552 is used to control the robot to interact in a multimodal manner that matches the target personality in response to an interaction operation for the robot, and the multimodal manner includes at least two of the following: voice, expression, and action; update module 5553 is used to update the target personality based on interaction feedback data during the robot interaction process, and control the robot to interact in a multimodal manner that matches the updated target personality.
[0147] In some embodiments, the configuration module is also used to display personality editing prompt information in response to a trigger operation on a personality setting entry associated with the robot; in response to target content edited based on the personality editing prompt information, extract personality parameters from the target content, and configure the personality of the robot to the target personality corresponding to the personality parameters.
[0148] In some embodiments, before configuring the personality of the robot to the target personality corresponding to the personality parameters, the device also includes: a determination module for determining at least one preset personality associated with the robot; matching the personality parameters with each preset personality to obtain a matching degree between the personality parameters and each preset personality; and determining the preset personality whose matching degree exceeds a matching degree threshold as the target personality corresponding to the personality parameters.
[0149] In some embodiments, the configuration module is also used to determine the influence data that affects the personality generation in response to the trigger operation of the personality setting entrance associated with the robot by the target object, and the influence data includes at least one of the following: historical interaction data between the target object and the robot, the interaction scene between the target object and the robot, the object data of the target object, and the environmental information of the robot; based on the influence data, the personality of the robot is predicted to obtain a target personality, and the personality of the robot is configured to the target personality.
[0150] In some embodiments, the configuration module is also used to extract features from the influencing data to obtain key features corresponding to the influencing data, wherein the key features include at least one of the following: emotional features, scene features, and environmental features; based on the key features, personality parameters that affect personality generation are generated; and personality mapping is performed based on the personality parameters to obtain a target personality corresponding to the personality parameters.
[0151] In some embodiments, the interaction module is also used to determine the second content of the robot's reply to the first content when the interaction operation indicates that the target object's interaction content with respect to the robot is the first content; perform semantic recognition and emotion recognition on the second content to obtain the target semantics and target emotion corresponding to the second content; in the mapping relationship under the target personality, determine the target expression associated with the target semantics based on the first mapping relationship between expression and semantics, determine the target action associated with the target semantics based on the second mapping relationship between expression and action, and determine the target timbre corresponding to the target emotion based on the third mapping relationship between emotion and timbre; control the robot to output the second content in voice with the target timbre, and control the robot to synchronously execute the corresponding target expression and target action.
[0152] In some embodiments, the interactive operation is a shaking operation for the robot, and the interactive module is further used to control the robot to perform feedback animation for the shaking operation in a multimodal manner that matches the target personality when the robot is not in a standby state; when the robot is in a standby state and the shaking operation is not triggered on the robot within a target historical period, reset the shaking countdown for the robot, and when the shaking countdown returns to zero, control the robot to perform feedback animation for the shaking operation in a multimodal manner that matches the target personality.
[0153] In some embodiments, the interactive operation is a pressing operation on a function key of the robot. The interactive module is further used to control the robot to broadcast introduction information about the robot in a multimodal manner that matches the target personality when the robot is in a normal state, and display the introduction information in the robot's interactive interface; when the robot is in an abnormal state, control the robot to output prompt information for the abnormal state in a multimodal manner that matches the target personality, and the abnormal state includes at least one of the following: no signal, battery level is lower than a battery threshold, and membership expires.
[0154] In some embodiments, after the introduction information is displayed in the interactive interface of the robot, the device further includes: a binding prompt module for displaying the device serial number and graphic code of the robot in the interactive interface in response to a shaking operation on the robot; wherein the graphic code is used for the identification terminal to download the robot's application and bind the robot through the application when recognizing the graphic code.
[0155] In some embodiments, the update module is also used to collect interaction feedback data of the target object interacting with the robot, the interaction feedback data including at least one of the following: feedback expression, feedback timbre, feedback action, and the degree of recognition of the target personality of the robot; perform feature extraction on the interaction feedback data to obtain feedback features of the target object; and update the personality parameters of the target personality based on the feedback features.
[0156] In some embodiments, the update module is also used to determine a first scene in which the target object is located based on the feedback features, and to determine a second scene to which the target personality is adapted by default; in the event that there is a conflict between the first scene and the second scene, the target personality is updated based on the first scene so that the updated target personality matches the first scene.
[0157] In some embodiments, after controlling the robot to interact in a multimodal manner that matches the updated target personality, the update module is further used to control the target personality of the robot to return from matching the first scene to matching the second scene when the target object exits the first scene.
[0158] The present invention provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the robot interaction processing method described in the present invention.
[0159] The embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will be caused to execute the robot interaction processing method provided by the embodiment of the present application, for example, Figure 3 The interactive processing method of the robot is shown.
[0160] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or various devices including one or any combination of the above memories.
[0161] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0162] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).
[0163] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.
[0164] Through the embodiments of the present application, through the user's personality configuration operation, the robot can quickly switch to the target personality, meeting the user's personalized needs in different scenarios. For example, in social situations, the user may want the robot to be more extroverted and humorous. Through personality configuration operation, the robot can adjust its personality parameters to display the corresponding personality traits. Secondly, the robot interacts with the user in a multimodal manner that matches the target personality, including voice, expression, and movement, greatly enhancing the naturalness and vividness of the interaction. For example, an extroverted and humorous robot can attract users with cheerful voice intonation, rich expression, and lively movement, giving users a more authentic and natural communication experience. Finally, during the interaction process, the robot updates the target personality in real time based on user feedback data and adjusts its multimodal interaction methods, achieving adaptive optimization. This feedback-based dynamic adjustment mechanism enables the robot to continuously learn and adapt to user preferences, thereby better meeting user needs and improving the quality and effectiveness of interaction. For example, if a user expresses dissatisfaction with a robot's behavior, the robot can adjust its personality parameters and interaction methods based on the feedback to better meet the user's expectations. Through this dynamic configuration, multimodal interaction and feedback-based update mechanism, the solution can significantly improve the robot's interactive performance, providing users with a more personalized, natural and adaptable interactive experience, thereby promoting the widespread application of robotics technology in more fields.
[0165] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.
Claims
1. A robot interaction processing method, characterized in that: The method comprises: In response to a character configuration operation for a robot, configuring the character of the robot to a target character; In response to an interactive operation directed to the robot, controlling the robot to interact in a multimodal manner that matches the target personality, the multimodal manner comprising at least two of the following: voice, expression, and action; During the robot interaction process, the target personality is updated based on the interaction feedback data, and the robot is controlled to interact in a multimodal manner that matches the updated target personality.
2. The method according to claim 1, characterized in that The step of configuring the robot's personality to a target personality in response to a personality configuration operation on the robot includes: In response to a triggering operation on a character setting entry associated with the robot, outputting character editing prompt information; In response to the target content edited based on the personality editing prompt information, personality parameters are extracted from the target content, and the personality of the robot is configured as the target personality corresponding to the personality parameters.
3. The method according to claim 2, characterized in that Before configuring the robot's personality to the target personality corresponding to the personality parameter, the method further includes: determining at least one predetermined personality associated with the robot; Matching the personality parameter with each of the preset personalities to obtain a matching degree between the personality parameter and each of the preset personalities; The preset personality whose matching degree exceeds the matching degree threshold is determined as the target personality corresponding to the personality parameter.
4. The method according to claim 1, wherein The step of configuring the robot's personality to a target personality in response to a personality configuration operation on the robot includes: In response to a triggering operation of a character setting entry associated with a robot by a target object, determining influencing data affecting character generation, the influencing data comprising at least one of the following: historical interaction data between the target object and the robot, interaction scenarios between the target object and the robot, object data of the target object, and environmental information of the robot; The character of the robot is predicted based on the influence data to obtain a target character, and the character of the robot is configured to be the target character.
5. The method according to claim 4, characterized in that The predicting the robot's personality based on the influence data to obtain a target personality includes: Extracting features from the impact data to obtain key features corresponding to the impact data, wherein the key features include at least one of the following: emotional features, scene features, and environmental features; generating personality parameters that influence personality generation based on the key features; Personality mapping is performed based on the personality parameters to obtain a target personality corresponding to the personality parameters.
6. The method according to claim 1, characterized in that The controlling the robot to interact in a multimodal manner matching the target personality comprises: When the interaction operation indicates that the target object's interaction content with respect to the robot is a first content, determining a second content in which the robot replies with respect to the first content; Performing semantic recognition and emotion recognition on the second content to obtain target semantics and target emotion corresponding to the second content; In the mapping relationship under the target personality, a target expression associated with the target semantics is determined based on a first mapping relationship between expression and semantics, a target action associated with the target semantics is determined based on a second mapping relationship between expression and action, and a target timbre corresponding to the target emotion is determined based on a third mapping relationship between emotion and timbre; The robot is controlled to output the second content in the voice of the target timbre, and the robot is controlled to synchronously execute the corresponding target expression and the target action.
7. The method according to claim 1, characterized in that The interactive operation is a shaking operation directed to the robot, and controlling the robot to interact in a multimodal manner matching the target personality includes: When the robot is in a non-standby state, controlling the robot to perform a feedback animation for the shaking operation in a multimodal manner that matches the target personality; When the robot is in a standby state and the shaking operation is not triggered on the robot within a target historical period, the shaking countdown of the robot is reset, and when the shaking countdown returns to zero, the robot is controlled to perform feedback animation for the shaking operation in a multimodal manner that matches the target personality.
8. The method according to claim 1, characterized in that The interactive operation is a pressing operation on a function key of the robot, and the controlling the robot to interact in a multimodal manner matching the target personality includes: When the robot is in a normal state, controlling the robot to broadcast introduction information for the robot in a multimodal manner that matches the target personality, and displaying the introduction information on the robot's interactive interface; When the robot is in an abnormal state, the robot is controlled to output prompt information for the abnormal state in a multimodal manner matching the target personality, and the abnormal state includes at least one of the following: no signal, power level below a power threshold, and membership expiration.
9. The method according to claim 8, characterized in that After displaying the introduction information in the interactive interface of the robot, the method further includes: In response to a shaking operation on the robot, displaying a device serial number and a graphic code of the robot in the interactive interface; The graphic code is used for the identification terminal to download the application of the robot when it recognizes the graphic code, and to bind the robot through the application.
10. The method according to claim 1, characterized in that The updating of the target personality based on the interactive feedback data includes: Collecting interaction feedback data of a target object interacting with the robot, the interaction feedback data including at least one of the following: feedback expression, feedback timbre, feedback action, and a degree of recognition of the target character of the robot; Extracting features from the interactive feedback data to obtain feedback features of the target object; The personality parameters of the target personality are updated based on the feedback features.
11. The method according to claim 10, characterized in that The updating of the target personality based on the feedback feature includes: Determining a first scenario in which the target object is located based on the feedback features, and determining a second scenario to which the target personality is adapted by default; In the case that there is a conflict between the first scenario and the second scenario, the target personality is updated based on the first scenario so that the updated target personality matches the first scenario.
12. The method according to claim 11, characterized in that After controlling the robot to interact in a multimodal manner that matches the updated target personality, the method further includes: When the target object exits the first scene, the target character controlling the robot is restored from matching the first scene to matching the second scene.
13. A robot interaction processing device, characterized in that: The device comprises: a configuration module, configured to configure the character of the robot to a target character in response to a character configuration operation on the robot; an interaction module, configured to, in response to an interactive operation directed to the robot, control the robot to interact in a multimodal manner that matches the target personality, the multimodal manner comprising at least two of the following: voice, expression, and action; An updating module is used to update the target personality based on interaction feedback data during the robot interaction process, and control the robot to interact in a multimodal manner that matches the updated target personality.
14. An electronic device, characterized in that: include: a memory for storing computer-executable instructions or computer programs; A processor, configured to implement the robot interaction processing method according to any one of claims 1 to 12 when executing the computer executable instructions or computer program stored in the memory.
15. A computer-readable storage medium, characterized in that Computer executable instructions or computer programs are stored, and when the computer executable instructions or computer programs are executed by a processor, the robot interaction processing method according to any one of claims 1 to 12 is implemented.
16. A computer program product comprising a computer program or computer executable instructions, characterized in that When the computer program or computer executable instructions are executed by a processor, the robot interaction processing method according to any one of claims 1 to 12 is implemented.
Citation Information
Cited By
Interaction control method and device, electronic equipment and computer readable storage medium
CN121028654A
Interactive control method and device, electronic equipment and computer readable storage medium
CN121028654B
Interaction control method, system and equipment of robot and medium
CN121560266A