Data processing method and electronic equipment
By obtaining the interaction data between the user and the device to generate and update the output parameters of the virtual image, the problem of the single virtual character image is solved, flexible interaction with the user is achieved, and the interactive experience is improved.
Patent Information
- Application Number
- CN202510729383.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-31
- Publication Date
- 2025-09-05
AI Technical Summary
In the existing technology, the virtual character images using artificial intelligence technology are relatively fixed and single, and cannot achieve flexible interaction with users in the real world.
By responding to target trigger events, the interaction data between the target user and the electronic device is obtained, a virtual image that can interact with the user is generated, and the output parameters of the virtual image, including emotional characteristics and environmental characteristics, are updated based on the interaction data to reflect the interaction process between the user and the application.
It realizes flexible interaction between the virtual image and the user, improves the authenticity and experience of the interaction, and can dynamically adjust the appearance and behavior of the virtual image to adapt to the user's different emotions and environmental changes.
Smart Images

Figure CN120595946A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a data processing method and electronic device. Background Art
[0002] With the continuous development of artificial intelligence (AI), the use of virtual characters in AI is increasing. For example, in AI-assisted teaching scenarios, virtual characters are often used to simulate real-world teachers. However, the virtual characters currently generated by AI applications are relatively fixed and monotonous, and cannot achieve flexible interaction with real-world users. Summary of the Invention
[0003] In one aspect, the present application provides a data processing method, comprising:
[0004] In response to a target triggering event, obtaining target interaction data between a target user and the electronic device;
[0005] generating a target virtual image capable of interacting with the target user based on the target interaction data;
[0006] updating output parameters of the target virtual image based on the obtained target reference data, wherein the target reference data may or may not be from the target user, and the output parameters can reflect the interaction process between the target user and a target application of the electronic device, wherein the target application is an application capable of invoking at least one processing model to at least generate or update the target virtual image;
[0007] Among them, under different triggering events, the target interaction data obtained is different.
[0008] In one possible implementation, obtaining target interaction data between a target user and an electronic device includes at least one of the following:
[0009] Based on the type of the target trigger event, at least one of target input data input by the target user to the target application, application data of the target application, or obtained behavior feature data of the target user is used as the target interaction data;
[0010] Based on the type of the target triggering event, the target interaction data is acquired in a corresponding acquisition manner.
[0011] In yet another possible implementation, generating a target virtual image capable of interacting with the target user based on the target interaction data includes:
[0012] Identifying a domain category to which the target interaction data belongs and a user intent represented, wherein the domain category is determined based at least on an attribute of a target content in the target interaction data;
[0013] At least one processing model is called based on the domain category and the user intention to generate a target virtual image for interacting with the target user.
[0014] In yet another possible implementation, generating a target virtual image capable of interacting with the target user based on the target interaction data includes:
[0015] Invoking at least one processing model to generate and process the obtained behavior characteristic data of the target user to obtain a user virtual image that matches the behavior characteristic data; and
[0016] Based on the application data of the target application and / or the target input data inputted by the target user to the target application, a target character virtual image capable of interacting with the user virtual image in a target virtual scene is generated.
[0017] In yet another possible implementation, updating the output parameters of the target virtual image based on the obtained target reference data includes:
[0018] Identifying emotional characteristic data and / or environmental characteristic data of a target user during interaction with the target application;
[0019] At least one of the constituent elements, character positioning, posture and movement, emotional expression, and sound parameters of the target virtual image is updated based on the emotional feature data and / or the environmental feature data.
[0020] In yet another possible implementation, identifying the target user's emotional characteristic data and / or environmental characteristic data during the interaction with the target application includes at least one of the following:
[0021] Identifying emotional attributes and / or environmental attributes carried in target input data input by a target user to a target application and / or feedback data of a response result of the target application, and obtaining the emotional feature data and / or environmental feature data based on the emotional attributes and / or the environmental attributes;
[0022] Obtaining behavioral characteristic data of the target user during the interaction process, and identifying the behavioral characteristic data to obtain the emotional characteristic data;
[0023] Obtaining environmental change data of the spatial environment in which the electronic device is located, and identifying the environmental change data to obtain the environmental characteristic data;
[0024] and / or,
[0025] The updating of the output parameters of the target virtual image based on the obtained target reference data further includes:
[0026] updating output parameters of the target virtual image based on attribute change information of the interaction content between the target user and the target application;
[0027] The output parameters of the target virtual image are updated based on usage change information of the usage scenario or usage purpose of the target application.
[0028] In yet another possible implementation, when the target user includes multiple users, the data processing method further includes:
[0029] Grouping the multiple users based on the emotional feature data and / or knowledge graph data of the multiple users to obtain multiple user groups;
[0030] Corresponding guidance data is generated based on the tag data corresponding to each user group, so as to guide the user group to complete the target task based on the corresponding guidance data.
[0031] In another possible implementation, the data processing method further includes at least one of the following:
[0032] Regrouping based on changes in the emotional feature data and / or knowledge graph data of users in each user group;
[0033] Generate corresponding positive evaluation content based on the emotional resonance points of multiple users in the user group.
[0034] In another possible implementation, the data processing method further includes at least one of the following:
[0035] When it is determined based on the user's emotional characteristic data that the user belongs to a set type of user, anonymizing the user's interaction data and / or exchanging conversation data through a target communication channel between the user and a target avatar;
[0036] updating output parameters of the target avatar based on the emotional characteristic data of the plurality of users;
[0037] Generate corresponding target virtual images for different user groups, or update the output parameters of the target virtual images based on the group emotional characteristic data of the user groups.
[0038] In another aspect, the present application further provides an electronic device comprising at least one processor and at least one processing model capable of running on the processor, wherein the processing model can be called by a target application to perform at least one of the following:
[0039] In response to a target triggering event, obtaining target interaction data between a target user and the electronic device;
[0040] generating a target virtual image capable of interacting with the target user based on the target interaction data;
[0041] updating output parameters of the target virtual image based on the obtained target reference data, wherein the target reference data may or may not be from the target user, and the output parameters can reflect the interaction process between the target user and a target application of the electronic device, wherein the target application is an application capable of invoking at least one processing model to at least generate or update the target virtual image;
[0042] Among them, under different triggering events, the target interaction data obtained is different. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.
[0044] Figure 1 A flowchart of the data processing method provided in this application;
[0045] Figure 2 A flowchart of another data processing method provided by this application;
[0046] Figure 3 This is an example diagram of a target virtual image generated based on target interaction data in this application;
[0047] Figure 4 A flowchart of another data processing method provided by this application;
[0048] Figure 5 An example diagram showing how the target virtual image displayed by the target application in this application changes as the environmental feature data changes;
[0049] Figure 6 An example diagram showing a target virtual image displayed by the target application in this application being updated and changed according to the user's emotions;
[0050] Figure 7 and Figure 8 Two example diagrams are shown for updating the expression, action, and sound parameters of a target avatar based on the emotional feature data carried by the target user's target input data;
[0051] Figure 9A schematic diagram of an interface of the target application in this application is shown;
[0052] Figure 10 A schematic diagram of the structure of the electronic device provided in this application. DETAILED DESCRIPTION
[0053] The embodiments of the present application are described below in conjunction with the drawings in the embodiments of the present application. The terms used in the implementation methods of the present application are only used to explain the specific embodiments of the present application and are not intended to limit the present application. It is known to those skilled in the art that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0054] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0055] like Figure 1 , showing a flow chart of the data processing method provided by the present application. The method of this embodiment is applied to an electronic device. In the present application, the electronic device has a target application or is associated with a target application, such as the electronic device is a terminal deployed with a target application; the electronic device can also be a server deployed with a target application, so that the target user can establish a connection with the target application through a browser or a client of the target application; the electronic device can also be a server that can establish a communication connection with the target application deployed on the terminal or a service node in a service system such as a cloud platform, so that the user can interact with the electronic device through the target application on the user end.
[0056] The target application is an application that can call at least one processing model to perform task processing. In this application, the target application is an application that can call at least one processing model to at least perform the generation or update of a virtual image.
[0057] Of course, the target application can also invoke the processing model to generate response content, teaching plans, or other feedback information, without limitation. For example, the target application can be an intelligent teaching application, enabling intelligent teaching of a subject or field through the target application. Another example is a social application or a gaming application, where users can engage in conversations, interactions, or consultations with virtual avatars. Of course, the target application can have other possibilities, without limitation.
[0058] The method of this embodiment may include:
[0059] S101 , in response to a target triggering event, obtaining target interaction data between a target user and an electronic device.
[0060] The target trigger event is a set event that requires starting, activating or outputting the virtual image.
[0061] For example, a target trigger event can open a target application or enter a designated interface of the target application (such as a main interface or an interface that requires interaction with a virtual image such as a virtual character) for the user. At this time, a target virtual image is presented to the user to prompt the user's operation or interact with the user.
[0062] For another example, considering that after the target application is launched or enters certain interfaces, it may be necessary to present a virtual image under certain conditions in order to provide services to users in a more vivid and vivid manner, based on this, the target trigger event can be one or more of the following events: the user inputs a question, the user publishes or uploads a specific task, the user's emotional changes are detected to meet the conditions, the target usage mode is launched (such as entering the learning mode, online course mode or knowledge test mode, etc.), the target application's usage scenario characteristics change, and the target application's usage purpose changes. Among them, the target application's usage scenario characteristics or usage purpose changes may be changes in the intelligent functions implemented by the user using the target application, the teaching content to be learned, or the game scenes to be participated in, which may require the launch of a virtual image or the replacement of a new virtual image.
[0063] The target user is the user who currently uses the target application.
[0064] The target interaction data between the target user and the electronic device is the input data currently input by the target user into the target application or the image data, sound data, or video surveillance data of the target user collected by the electronic device under the program instructions of the target application. The target interaction data can reflect the target user's demand for the virtual image.
[0065] It can be understood that since the interaction between the target user and the target application is related to the interface state or operating mode of the target application itself, in this application, the target interaction data can not only include the target user's behavioral data, such as the target input data input by the user to the target application, and the target user's behavioral feature data; it can also include the application data of the target application, etc.
[0066] The target input data may include, but is not limited to, question data input by the target user to the target application, feedback data in response to the question, published task data, interactive operation data or configuration operation data acting on the target application, etc. The target input data may include data in multiple modalities, such as image data, text data, or voice data, without limitation.
[0067] The target user's behavioral characteristic data may include: facial expression characteristics (such as frowning, laughing, or glaring), body posture characteristics (such as supporting the chin, nodding, shaking the head, raising the hand, or maintaining a certain posture for a certain period of time), voice characteristics (such as voiceprint characteristics or voice tone characteristics), and emotional characteristics, etc., which are related to the user's own behavior. Among them, emotional characteristics can be analyzed based on the user's behavior and can reflect the characteristics or behavioral manifestations of the user's emotional category. For example, emotional characteristics can include silence, laughter, crying, depression, or mental focus.
[0068] The application data of the target application may be data currently possessed by the target application, which can reflect the user interaction features or interaction behaviors currently supported by the target application. For example, the application data of the target application may be at least part of the data such as the application type of the target application, application identification, application interface data, task data executed by the target application, and usage scenario data of the target application. The application data of the target application may also reflect the characteristics of the virtual image that is suitable for presentation or the characteristics of the virtual image that the user may expect to see. For example, if the target application is in an application mode such as English test, mathematics teaching or physics training, a virtual image for simulating an invigilator, training lecturer or teaching teacher is presented.
[0069] In this application, the target interaction data to be obtained may be one or more of the above situations according to actual needs, without specific limitation.
[0070] The target interaction data obtained under different triggering events are different. For example, when the application interface of the target application is opened, the subject category to which the application interface belongs or the application type to which the target application belongs can be used as the target interaction data.
[0071] Of course, the method of obtaining target interaction data will be different under different trigger events.
[0072] For example, in this application, the target interaction data may be obtained in at least one of the following two ways:
[0073] In one possible scenario, based on the type of the target trigger event, at least one of target input data input by the target user to the target application, application data of the target application, or obtained behavior characteristic data of the target user may be used as the target interaction data.
[0074] The type of target trigger event can represent the generation method or event category of the target trigger event. For example, the types of target trigger events can be divided into target application opening, target usage mode startup or switching (such as entering learning mode, homework evaluation mode or test mode, etc., or switching from learning mode to test mode, etc.), user emotion change, question asking, task release (such as submitting homework or releasing task requirements to be executed, etc.), or usage scenario change (such as switching learning subjects or switching game types in a certain usage mode; or switching from learning usage scenario to leisure scenario).
[0075] It is understood that the target interaction data that requires attention varies depending on the type of target trigger event. Therefore, based on the type of target trigger event, target interaction data related to the type of target trigger event can be obtained, and the specific data can be set according to actual needs. For example, if the target trigger event type is a change in user emotion, then one or both of the target user's behavioral characteristic data and target input data can be obtained when the target trigger event occurs. For another example, if the target trigger event is the activation of a target usage mode, then at least the application data of the target application and the target input data of the target user can be obtained.
[0076] In another possible scenario, the target interaction data is acquired in a corresponding acquisition manner based on the type of the target triggering event.
[0077] In this possible scenario, different target triggering event types may result in different methods for obtaining target interaction data, and accordingly, the target interaction data obtained may also vary. For example, at least one acquisition method suitable for different types of triggering events may be pre-configured, so that after obtaining a target triggering event, the target interaction data can be obtained using the corresponding acquisition method based on the type of the target triggering event.
[0078] The method for obtaining the target interaction data can be reflected in the sensor or collector used to obtain the target interaction data and the source of the data. The sensor or collector may include but is not limited to: a camera, a gravity sensor, a temperature sensor, and an audio collector. The source of the data may be obtaining interface input data, retrieving operation data, or background data.
[0079] For example, when the target trigger event is an emotion change, the camera can be used to capture facial and body images of the user, and input data such as the voice signal collected by the audio collector and the context of the target user input obtained through the application interface of the target application can also be obtained.
[0080] S102: Generate a target virtual image capable of interacting with a target user based on the target interaction data.
[0081] The target avatar may be a virtual person or cartoon character created based on computer graphics and artificial intelligence (AI) technology. For example, the target avatar may be an AI digital human simulating a teacher, coach, lawyer, salesperson, or conversational user.
[0082] The interaction between the target user and the target virtual image can simulate the interaction effect between the target user and the real user.
[0083] Among them, based on the different target interaction data, the role represented by the generated target virtual image and the attribute characteristics such as the appearance, emotion, clothing, action, age and gender of the target virtual image will also be different.
[0084] S103: Update the output parameters of the target virtual image based on the obtained target reference data.
[0085] The target reference data may or may not come from the target user. For example, the target reference data can be the emotions represented by the target user's behavior, the emotions that can be triggered in others, or the target user's behavioral characteristics. It can also be the target user's environment or environmental data detected by electronic devices. The environmental data can include weather data, seasonal characteristic data, and location data representing indoor or outdoor locations.
[0086] The output parameters of the target avatar can reflect the interaction process between the target user and the target application of the electronic device. For example, the output parameters of the target avatar can reflect the target user's emotional state during the interaction with the target application of the electronic device, the content of the interaction (such as the voice and text output of the simulated target avatar, as well as the response feedback to the user's input questions), and the smoothness of the interaction.
[0087] As can be seen from the foregoing, the target application is an application that can call at least one processing model to at least generate or update the target virtual image.
[0088] The output parameters of the target avatar are used to characterize at least the target avatar's performance characteristics. These performance characteristics are not limited to visual features, but also include audio features of the target avatar's output. For example, the output parameters of the target avatar may include one or more of the following: the target avatar's constituent elements, character, posture, movements, emotional expressions, clothing, and output statements associated with the target avatar. Of course, they may also include the text or voice output of the target avatar, allowing users to intuitively experience text or voice interaction with the target avatar.
[0089] As can be seen from the above, in this application, after obtaining the target interaction data between the target user and the electronic device in response to the target trigger event, a target virtual image capable of interacting with the target user will be generated based on the target interaction data, thereby being able to combine the interaction data between the target user and the electronic device to present the virtual image more reasonably and flexibly. On this basis, after generating the target virtual image, this application will also update the output parameters of the target virtual image based on the obtained target parameter data. Since the output parameters can reflect the interaction process between the target user and the target application of the electronic device, the target virtual image can provide more realistic and vivid interaction feedback, thereby improving the interaction flexibility and interaction experience.
[0090] In the present application, there can be one or more target virtual images generated, and there can be multiple specific generation methods, which are not specifically limited.
[0091] In one possible scenario, taking the example where the target interaction data includes the behavioral characteristic data of the target user and also includes at least one of the application data of the target application and the target input data, generating the target virtual image based on the target interaction data can be: calling at least one processing model to generate and process the obtained behavioral characteristic data of the target user to obtain a user virtual image that matches the behavioral characteristic data; and, based on the application data of the target application and / or the target input data input by the target user to the target application, generating a target character virtual image that can interact with the user virtual image in the target virtual scene.
[0092] Among them, the processing model can be any model based on artificial intelligence technology, such as a multimodal large language model, a raw image model, a 3D modeling model, a 3D animation model, or a processing module combined with artificial intelligence technology.
[0093] In this possible scenario, the user avatar generated based on the target user's behavioral characteristic data can be an avatar that can present or exhibit the target user's role in the interaction and the corresponding behavioral characteristics, so as to simulate the target user's actual behavioral characteristics through this avatar. For example, if the target user's behavioral characteristics indicate that the target user is a student participating in a study, and the target user is currently in a state of raising their hand and frowning, then a avatar representing the student can be generated with their hand raised and frowning. It can be seen that this user avatar is essentially generated for the target user and represents the target user.
[0094] In this possible scenario, in addition to generating a user avatar for the target user, a target character avatar is also generated that can interact with the user avatar. The target character avatar is generated based on at least one of the target application's application data and the target user's input data into the target application, and is a avatar character that can respond appropriately to the target user's behavioral characteristic data.
[0095] For example, the application data and user input data of the target application represent the target user's learning of the subject of physics through the target application. In this case, a target role virtual image representing the teacher (physics teacher role) can be generated to simulate real interaction through the interaction between the target role virtual image and the user virtual image (student role), thereby improving the sense of immersion and participation in the interaction.
[0096] The target virtual scene is the virtual scene in which the user's avatar and the target avatar reside. The type, background, virtual location, and virtual environment of the target virtual scene are all related to the user's avatar and the target avatar, as well as the application data of the target application and the target input data of the target user. For example, the target virtual scene can be a teaching scene (e.g., a virtual indoor classroom scene or an outdoor teaching scene), an examination scene, a sports training scene, a sales location scene, or a gaming scene.
[0097] Specifically, the specific interaction between the user's virtual image and the target virtual image may vary depending on the application data of the target application and / or the user's input data.
[0098] For example, assuming that based on the application data or target input data of the target application, it is determined that the target user is currently studying mathematics based on the target application, then a student virtual image corresponding to the target user can be constructed, and a teacher virtual image can be constructed, and a virtual classroom scene can be generated in which the student virtual image interacts with the teacher virtual image. The virtual classroom scene can include a blackboard and desks as the background.
[0099] Of course, the interaction between the user's virtual image and the target virtual image in the target virtual scene can also be knowledge questions and answers, content explanations, or sports training, etc.
[0100] In this application, when generating a target virtual image based on target interaction data, there are many possibilities for analyzing the target interaction data and finally generating the target virtual image. The following is an example of an implementation method. Figure 2 , shows another flow chart of the data processing method provided by the present application. The method of this embodiment may include:
[0101] S201 , in response to a target triggering event, obtaining target interaction data between a target user and an electronic device.
[0102] S202: Identify the domain category to which the target interaction data belongs and the user intention represented.
[0103] The domain categories of the target interaction data may include different subject categories, industry fields, or application fields. Subject fields may include language, history, geography, physics, or chemistry. Industry fields may include architecture, civil engineering, economics, new energy, or biomedicine. Application fields may include industry or agriculture.
[0104] The domain category to which the target interaction data belongs is determined based at least on the attributes of the target content in the target interaction data.
[0105] The target content may be content in the target interaction data that reflects the domain information to which the target user may interact with the target application. The target content may vary depending on the target interaction data. For example, when the target interaction data includes user input data, the target content may be the problem description information in the user input data, or keyword data related to the domain. When the target interaction data includes application data of the target application, the target content may be the interface data of the target application or the usage scenario corresponding to the application interface.
[0106] The attributes of the target content may be category attributes or semantic attributes of the target content, and may also include sentiment attributes and the like.
[0107] User intent is the intention represented by the target interaction data, that is, the purpose the target user hopes to achieve through the target application. For example, user intent can include learning subject content, calculation, explaining concepts, grading homework, obtaining industry reports, obtaining personalized images, obtaining personalized text descriptions, or generating application code. For example, if the user intent is calculation, it means that the user wants to obtain solution ideas and process related to calculations.
[0108] In particular, under the premise that the target interaction data includes target input data, the complexity of processing the target input data can also be represented by the user intention.
[0109] S203: Call at least one processing model based on the domain category and the user intention to generate a target virtual image for interacting with the target user.
[0110] It can be understood that since the domain category and user intent can reflect the disciplines and fields that the target user needs to interact with, and can characterize the specific purpose of the target user's use of the target application, etc., based on the domain category and user intent, the role, appearance characteristics, expressions, actions, output text, output voice, etc. of the target virtual image can be determined.
[0111] Combine Figure 3 For example, in Figure 3 In the example, it is assumed that the domain category determined based on the target interaction data is mathematics, and the user's intention is to learn mathematics courses. On this basis, this application will generate a virtual character representing the teacher, such as Figure 3 A smiling and amiable virtual female teacher image was generated.
[0112] As can be seen, this application can dynamically and reasonably adjust the appearance, movements, and expressions of the target virtual avatar based on the teaching content, interactive content characteristics, and methods of different subjects or fields. For example, a math teacher may focus more on logical rigor, while a Chinese teacher may focus more on emotional expression and literary style.
[0113] In this embodiment, the specific implementation of calling the processing model to generate the target virtual image that interacts with the target user can be any of the aforementioned ones, without specific limitation.
[0114] For example, at least one processing model can be called to generate and process the obtained behavioral characteristic data of the target user to obtain a user virtual image that matches the behavioral characteristic data; and based on the domain category and user intention, a target character virtual image that can interact with the user virtual image in the target virtual scene is generated.
[0115] S204: Update output parameters of the target virtual image based on the obtained target reference data.
[0116] The target reference data may or may not be from the target user. The output parameter can reflect the interaction process between the target user and the target application of the electronic device, where the target application is an application capable of invoking at least one processing model to at least generate or update the target virtual image.
[0117] In any of the above embodiments of the present application, there are also multiple implementations for updating the output parameters of the target virtual image based on the obtained target reference data. The following is an example of an implementation. Figure 4 , shows another flow chart of the data processing method provided by the present application. The method of this embodiment may include:
[0118] S401 , in response to a target triggering event, obtaining target interaction data between a target user and an electronic device.
[0119] S402: Generate a target virtual image capable of interacting with a target user based on the target interaction data.
[0120] The above two steps can be referred to the relevant introduction of any of the previous embodiments and will not be repeated here.
[0121] S403: Identify emotional characteristic data and / or environmental characteristic data of the target user during the interaction process with the target application.
[0122] In this embodiment, the target reference data includes one of emotional feature data and environmental feature data during the interaction between the target user and the target application.
[0123] Among them, the emotional feature data is used to characterize the target user's emotions, expressions, personality state, physical state or psychological state, etc. For example, the emotional feature data can reflect whether the user's mood is relatively low or whether the psychological pressure is too great. The emotional feature data can be determined based on one or more of the following: the facial image, body image, sound information of the target user, the data content input into the target application (such as questions raised by the user, the answers or reply information given by the user to the questions given by the user to the target virtual image or target application), and the progress status of the task executed by the target application (such as learning progress, teaching progress, learning acceptance, teaching progress, learning difficulty feedback). For example, if the input content input by the target user is "This question is so difficult, I can't understand it at all, it's so uncomfortable", it can be determined that the target user is currently in a low mood and has great psychological pressure.
[0124] The environmental characteristic data may be one of the data such as the ambient temperature, weather, season, ambient space, ambient brightness and ambient noise of the environment where the target user is located or the environment mentioned during the interaction between the target user and the target application.
[0125] There are many possibilities for the specific implementation of obtaining environmental characteristic data and / or emotional characteristic data in this application, and there are no specific limitations.
[0126] For example, the corresponding environmental characteristic data and / or emotional characteristic data may be obtained through at least one of the following possible situations:
[0127] In one possible scenario, the emotional attributes and / or environmental attributes carried in the target input data input by the target user to the target application and / or the feedback data of the response result to the target application are identified, and emotional feature data and / or environmental feature data are obtained based on the emotional attributes and / or the environmental attributes.
[0128] Among them, after the target virtual image is generated, the target input data input by the target user to the target application can be different from the target input data obtained in response to the target trigger event. For example, it can be input data obtained at different times, or it can be feedback data of the inference results obtained after the target user calls other processing models for the target application to infer the questions asked by the target user or the tasks issued, such as evaluation data or further supplementary input data of the question answering data or the image generation results corresponding to the image generation task. Of course, in this application, in order to be able to dynamically update the target virtual image in real time, the target input data input by the target user will be continuously obtained.
[0129] The data information that the target input data here can contain is the same as the target input data mentioned above and will not be repeated here.
[0130] Among them, the response result of the target application is the response content output by the target application or the display interface switched to, and other information. The feedback data can be the feedback given by the target user in response to the response result of the target application. Based on the feedback data, it can reflect the user's satisfaction with the response result of the target application, the degree of acceptance, and the user's feelings about the difficulty of teaching, and other emotional attributes. It can also reflect the environmental attribute characteristics that the user wants to see. For example, the feedback data can be the satisfaction or satisfaction of the target user's choice. If the user is relatively satisfied, it means that the user is relatively happy or cheerful. For another example, if the feedback data expresses the desire to switch to the environment where the target virtual image is located, then relevant environmental data can be extracted from the feedback data.
[0131] It is understandable that since the target input data can include the input content of the target user, or the operations performed by the target user on the target application, the target user's emotions, learning status and other emotional information can be reflected in combination with the input content or operations corresponding to the target input data. Of course, it is also possible to extract environment-related attribute information from the interface data presented by the target application triggered by the input content or operation. For example, if the user inputs "Please introduce the relevant knowledge about Antarctica", it can be determined that the environmental characteristics corresponding to the environment that requires interaction are Antarctica, which has a relatively cold climate.
[0132] In particular, the electronic device can be configured with a user vocabulary for characterizing different emotional characteristics of the user. The user vocabulary can include emotional words corresponding to multiple different emotional types. On this basis, based on the emotional words in the target input data or feedback data, the emotional characteristics expressed by the target input data or feedback data can be determined.
[0133] The user vocabulary can be configured based on the type of target application and actual needs, without any specific restrictions. For example, if the target application is an auxiliary teaching application, the user vocabulary can include emotional words corresponding to the following categories of emotions:
[0134] Knowledge comprehension categories include: difficult to understand, confusing, and bewildered. For example, if the user input data includes "This part of knowledge is difficult to understand," it can help determine that the emotion expressed by the user input data is difficult to understand knowledge.
[0135] Learning interest categories include: like, interest, and boredom; for example, if the user input data includes "this class is so boring", it can be determined that the user has the emotional characteristics of interest.
[0136] The learning stress category includes anxiety, tension, and high pressure. For example, if the user input data includes "The exam is coming soon and I am very anxious", it can indicate that the user has the emotional characteristics of learning stress.
[0137] The learning achievement category includes pride, sense of accomplishment, and frustration. For example, if the user input data includes "I feel a sense of accomplishment after solving this difficult problem", it indicates that the user has a positive emotion of learning achievement.
[0138] The learning attitude category includes: coverage, willingness, positive and negative, etc.; for example, user input data includes "I am willing to actively participate in group discussions", indicating that the user has a positive learning attitude.
[0139] It is understandable that the application data of the target application can also reflect the environmental characteristic data during the interaction process. For example, the contextual information during the interaction process presented in the application interface of the target application can include environmental-related characteristic data. For example, the application interface of the target application is in the interface of geographic information explanation, and the content of the explanation will involve areas with different environmental characteristics. In order to achieve the effect of immersive learning, the environmental characteristic data can be determined based on the application data of the target application. On this basis, it can also be: identifying the target input data input by the target user to the target application, the application data of the target application, and / or the emotional attributes and / or environmental attributes carried in the feedback data of the response result of the target application, and obtaining the emotional characteristic data and / or environmental characteristic data based on the emotional attributes and / or the environmental attributes.
[0140] Of course, since the target input data can reflect the context information presented by the application interface of the target application, the environmental feature data contained in the content of the application interface can also be directly or indirectly represented based on the target input data.
[0141] For easier understanding, see Figure 5 , Figure 5 An example diagram of an interaction interface between a target user and a target application is shown.
[0142] exist Figure 5 It can be seen that the target user has started the natural environment explanation course through the target input operation. Since the geographical area where the target application starts to explain is Antarctica, the target virtual image is a virtual character in Antarctica (such as Figure 5 (As shown in the image of the first virtual character in the video), the virtual character is wearing a padded coat to match the Antarctic environment.
[0143] On this basis, if the target user inputs "Teacher, please tell me about Sanya", the target application will switch to the corresponding geographical knowledge explanation of Sanya, and will update the background and clothes of the virtual character, such as Figure 5 As shown in the second image of the virtual character, when explaining the geographical knowledge related to Sanya, the virtual character's clothes were changed into relatively refreshing thin clothes, and the background included the seaside and coconut trees.
[0144] In another possible scenario, behavioral characteristic data of the target user during the interaction process is obtained, and the behavioral characteristic data is identified to obtain emotional characteristic data.
[0145] It is understandable that the target user's behavioral characteristic data may be features such as facial expressions, body postures, and voice tones, and these behavioral characteristic data can reflect the user's current emotional characteristics.
[0146] For example, taking the scenario of assisted teaching based on the target application as an example, by obtaining the target user's facial expressions, the voice of answering questions, and body posture during the learning process with the help of the target application, the target user's learning enthusiasm, concentration, and emotional state can be determined.
[0147] In another possible scenario, environmental change data of the spatial environment in which the electronic device is located is obtained, and the environmental change data is identified to obtain environmental characteristic data.
[0148] For example, environmental information corresponding to the target user's location is obtained through a weather monitoring agency or a weather monitoring application; or, with the target user's permission, environmental data is determined by obtaining an environmental image of the target user's environment.
[0149] S404: Based on the emotional feature data and / or environmental feature data, update at least one of the constituent elements, character positioning, posture and movement, emotional expression, and sound parameters of the target virtual image.
[0150] Among them, the constituent elements can be clothing, props, hair accessories, decorations, etc., which can highlight the character characteristics of the virtual image.
[0151] The role positioning can be the role type of the avatar. For example, the avatar can represent a subject teacher, trainer, coach, assistant, playmate, doctor, or lawyer.
[0152] Posture actions can be the movement characteristics and body postures of the avatar. For example, the target avatar's posture actions can include but are not limited to giving a thumbs-up, resting one's chin on one's hand while thinking, adjusting one's glasses, scratching one's head, or waving one's hand.
[0153] Emotional expressions may include, but are not limited to, expressions of encouragement, praise, or expressions of heat or comfort that change with ambient temperature.
[0154] Voice parameters are the sound parameters of the output sound configured for the target avatar. These include speech rate, intonation, volume, sound content (e.g., words), and the emotional characteristics that the sound is intended to convey. As the emotions represented by the emotional characteristics data change, the voice parameters can be modified to update the target avatar's voice style. For example, a stern tone can be changed to a gentle one, or the content of the sound can be changed from critical to encouraging.
[0155] For example, during the knowledge explanation process, when the target user is in a bad mood, the voice output by the target virtual image needs to be relatively gentle, including content such as "Persevere, don't give up". When the target user's concentration is not high, the target virtual image can use a harsher tone to output critical sentences such as "Please concentrate on listening and don't look elsewhere".
[0156] Among them, in order to make the sound content in the sound parameters output by the target virtual image more reasonable, the present application can also construct a vocabulary library that can be output by the target virtual image. The vocabulary library can include at least one word corresponding to each of multiple emotion types for expressing different emotions, without any specific restrictions.
[0157] It can be understood that, while being based on emotional feature data and / or environmental feature data, the target interaction data of the target user can also be combined to update at least one of the constituent elements, character positioning, posture movements, emotional expressions and sound parameters of the target virtual image, so as to more reasonably update the output parameters of the target virtual image.
[0158] In this embodiment, after generating the target virtual image based on the target interaction data, the present application will also identify at least one of the emotional characteristic data and environmental characteristic data of the target user during the interaction with the target application, and based on at least one of the emotional characteristic data and environmental characteristic data, update the constituent elements, role positioning and emotional expressions of the target virtual image, so that the performance characteristics of the target virtual image can change with at least one of the emotional characteristics and environmental characteristics of the target user during the interaction with the target application, thereby improving the target user's sense of immersion in the interaction with the target application and improving the interactive experience.
[0159] For example, when updating the target virtual image based on environmental feature data, by perceiving external environmental factors such as the season and weather in which the target user is located, as well as environmental information expressed by the current teaching content of the target application, the appearance and clothing of the target virtual image such as the AI digital human can be dynamically adjusted to improve the fit with the teaching scene or other current applicable scenes of the target application, making the learning or task interaction experience more realistic.
[0160] For example, in terms of updating the target virtual image based on emotional feature data, since the emotional feature data is determined based on the target user's interaction status, interaction enthusiasm (such as whether actively participates in discussions), feedback (such as learning difficulty), and learning (or task) stage, learning (or task) progress, knowledge mastery, target user's facial image, target user's input voice and other modal data during the process of using the target application to learn or perform other tasks, therefore, updating the target virtual image in combination with the emotional feature data can achieve reasonable adjustment of the target virtual image based on multiple aspects such as the user's feedback on learning (or task) progress, acceptance level and learning difficulty.
[0161] For example, taking auxiliary teaching based on target applications as an example, for students with slower progress, the target virtual image such as AI digital people will automatically reduce the difficulty of the learning content and increase auxiliary interaction; it can also identify the emotional state of the students (such as anxiety, confusion, excitement, etc.), and dynamically adjust the expression and body language of the AI digital people, and provide appropriate emotional feedback to enhance the students' emotional resonance. For example, by analyzing the information of multiple modalities such as the student's input voice, facial image, input questions, and learning progress, it is found that the student's voice is trembling and the speed of speech is accelerated, and the text contains the word "very nervous". It can be comprehensively judged that the student is in a nervous state. The simulated virtual teacher can be adjusted to output soothing and encouraging words, and present a smiling expression, etc., to give appropriate interactive feedback. In order to facilitate understanding of the benefits of this embodiment, the following examples are combined with several application scenarios for explanation.
[0162] like Figure 6, which shows an example diagram of a target virtual image displayed by a target application being updated as the user's emotions change.
[0163] exist Figure 6 In this example, the target user learns mathematics through the target application.
[0164] again Figure 6 It can be seen that after the target user inputs a question into the target application, the electronic device can determine that the question is about the subject of mathematics 601, and will generate a virtual character image 602 representing a mathematics teacher. The background of the virtual character image is the scene of the teacher's mathematics classroom. At the same time, the response information 603 to the question will be output on the interface of the target application.
[0165] After the target user enters input data 604 in response to the reply question 603, the electronic device can determine from this input data that the target user's emotional characteristics indicate that the target user is currently in a state of confusion and anxiety. Based on this, the electronic device will synchronously adjust the movements of the avatar representing the math teacher to produce an avatar image 605. The avatar image will appear to be in a state of contemplation with its chin in its hand, and a reply message 606 will be output on the interface of the target application.
[0166] After the target user sees the reply message 606 confirming that the problem has been solved, the target user enters input data 607, and the input data 607 includes "Wow, teacher, I understand." It is obvious that the target user is in high spirits and presents a positive emotional state. Then the hand movement of the virtual character image will be updated to show a thumbs-up action, and the virtual character image 608 will be obtained, and the corresponding reply content 609 will be output at the same time.
[0167] The following combination Figure 3 、 Figure 7 and Figure 8 , the situation of updating the expression, action and sound parameters of the target virtual image based on the emotional feature data carried by the target user's target input data is explained.
[0168] exist Figure 3 The target user learns mathematics courses through the target application, and the target virtual image representing the teacher (hereinafter referred to as the virtual teacher) gives a mathematics problem for the target user as a student to solve.
[0169] exist Figure 3 Based on Figure 7It can be seen that the answer input by the target user is correct. Combining the math problem and the question-answer response, it can be determined that the target user's answer can represent that the student has a serious and positive learning attitude. On this basis, the virtual teacher will give positive and affirmative feedback. Figure 7 , the virtual teacher will smile and applaud with his hands, and will also output a voice message containing "The classroom speech is so wonderful", and the same text content as the voice message will be output synchronously on the application interface.
[0170] exist Figure 3 Based on Figure 8 It can be seen that the target user's answer to the math problem given by the virtual teacher is wrong, and it is an obvious calculation error. Therefore, combined with the answer, it can be analyzed that the target user has emotional characteristics such as a careless learning attitude. Therefore, the virtual teacher will show a frowning expression and will simultaneously output the voice "You are too careless when doing the problem, you must carefully review the problem" in a relatively harsh tone.
[0171] It is understandable that in Figure 4 In the embodiment, an implementation method of updating the output parameters of the target virtual image is used as an example to illustrate. In actual applications, in addition to being based on at least one of the above-mentioned emotional feature data and environmental feature data, the output parameters of the target virtual image can also be updated based on other interactive information between the target user and the target application or part or all of the information in the application data of the target application.
[0172] For example, in a possible implementation, updating the output parameters of the target virtual image based on the obtained target reference data may further include at least one of the following two situations:
[0173] In one possible scenario, the output parameters of the target virtual image may be updated based on the attribute change information of the interaction content between the target user and the target application.
[0174] The interactive content may include one or both of the content output by the target application and the target input data input by the target user.
[0175] The attributes of interactive content can include the category, chapter (e.g., the learned subject chapter), associated author, and application usage characteristics (e.g., learning attitude) reflected by the interactive content. For example, in the scenario of implementing assisted teaching based on the target application, attribute change information of interactive content can include changes in the subject to which the interactive content belongs, changes in chapters within the subject, etc.
[0176] Among them, for the update of the output parameters of the target virtual image, please refer to the previous related introduction, which will not be repeated here.
[0177] Taking the scenario of implementing assisted teaching based on the target application as an example, if the interaction content between the target user and the target application switches from explaining mathematical formulas to analyzing poetry, the role positioning, interaction method, and expression style of the target virtual image can be adjusted.
[0178] In another possible scenario, the output parameters of the target virtual image may be updated based on usage change information of the usage scenario or usage purpose of the target application.
[0179] Among them, the usage scenario or usage purpose of the target application can refer to the previous introduction. When the usage scenario and usage purpose change, the output parameters of the target virtual image can be updated to output parameters that match the changed usage scenario or usage purpose. For example, still taking the teaching scenario as an example, the interactive style, expression or movement of the target virtual image can be adjusted synchronously according to changes in the teaching links provided by the target application (such as homework correction or question answering, etc.), or changes in teaching strategies.
[0180] It is understandable that in the above embodiments, the target user is taken as an example. In actual applications, the target users of the target application may include multiple users, and different users may log in to the target application through different clients and establish connections with electronic devices. In this case, in order to be able to provide more targeted services to each user through the target application, in this application, the multiple users can also be grouped based on the emotional feature data and / or knowledge graph data of the multiple users to obtain multiple user groups. On this basis, corresponding guidance data can be generated based on the label data corresponding to each user group to guide the user group to complete the target task based on the corresponding guidance data.
[0181] The user's emotional characteristic data may be emotional characteristic data of the user determined within a recently set period of time, such as emotional characteristic data determined based on interaction data between the user and the target application within the set period of time. The acquisition of emotional characteristic data and the specific data content can be found in the previous related introduction and will not be repeated here.
[0182] The user's knowledge graph data can reflect the target user's identity attributes and the target user's usage characteristics of the target application. For example, the user's knowledge graph data can indicate the user's identity, historical relationships with other users, learning progress based on the target application (such as the current chapter being learned), and knowledge mastery.
[0183] It is understandable that based on the emotional characteristic data and / or knowledge graph data of multiple users, users with similar emotional characteristics and similar usage characteristics of the target application (such as learning progress or knowledge mastery) can be divided into the same user group.
[0184] For example, based on Euclidean distance calculation, users with similar sentiment feature data or knowledge graph data are identified and assigned to the same user group.
[0185] For another example, the present application can also quantify the emotional characteristic data of each user to obtain an emotional state score for each user, for example, the emotional state score is 1-10, and users with the same emotional state score are grouped into user groups.
[0186] The user group label data is used to distinguish or locate user groups. For example, the user group label data can be a user group ID, or identification data that can characterize the emotional characteristics or usage characteristics of the user group. For example, if the target application is an application for auxiliary teaching, label data can be generated for each user group based on the student's learning enthusiasm, learning progress, or knowledge mastery. For example, the user group label data can be divided into an enthusiastic group, a top student group, and a lagging group.
[0187] The label data of the user group can be determined based on the emotional feature data and / or knowledge graph data of each user in the user group. For example, the label data of the user group can represent that each user in the user group has common features.
[0188] The guidance data is data used to cause the user group to perform task operations based on the target application. For example, the guidance data may include learning strategies and teaching strategies for guiding users in the user group to learn, or teaching content, teaching methods, and incentive mechanisms suitable for the user group.
[0189] Among them, the encouragement mechanism can be an encouragement method or encouragement language. In this application, an encouragement vocabulary library can be configured in the electronic device, and different encouragement words are configured for different types of user groups.
[0190] For example, the encouragement word library may include encouragement words corresponding to the following types of encouragement:
[0191] Positive encouragement: Come on, keep going, don’t give up, try again;
[0192] Passive soothing: I know you’re sad, I understand your frustration, don’t worry, take your time;
[0193] Question-guided questions: Let’s find the answer together. Where do you plan to start solving this problem?
[0194] Neutral guidance: Complete this assignment, read this article, do this experiment, memorize these words, organize your notes;
[0195] Praise: You have done an excellent job, you have done it perfectly, your results are amazing, this performance is exemplary, the quality of the work you submitted is very high.
[0196] It is understood that when the target user includes multiple users, grouping the multiple users can be performed periodically, or when the emotional feature data and / or knowledge graph data of the multiple users changes, or can be performed irregularly, without limitation. It can be seen that grouping the multiple users can be performed before or after obtaining the target trigger event, or before or after obtaining the target parameter data, without limitation.
[0197] Furthermore, in the present application, regrouping can also be performed based on the change information of the emotional characteristic data and / or knowledge graph data of users in each user group. Among them, regrouping can also be performed periodically according to a fixed time period (for example, regrouping is performed once every two weeks). In this case, when the moment of regrouping is reached according to the set time period, it is necessary to recalculate or determine the emotional characteristic data and / or knowledge graph data of each user in the user group. Of course, regrouping can also be performed when it is detected that the data with changes in the emotional characteristic data and / or knowledge graph data of users in the user group exceeds a set ratio, and there is no specific restriction.
[0198] In one possible scenario, when multiple users are divided into multiple user groups, the present application can also generate corresponding positive evaluation content based on the emotional resonance points of multiple users in the user group.
[0199] Among them, the emotional resonance points of multiple users in a user group can be the same emotional characteristics expressed by multiple users. For example, taking the teaching scenario as an example, in a history class discussion, if a student tells a historical story and other students in the user group to which the student belongs are captured showing focused eyes, nodding, collective applause and other body movements or positive voice feedback, etc., then it can be determined that there are emotional resonance points among multiple users in the user group.
[0200] Based on the emotional resonance types of multiple users' emotional resonance points, this application can generate corresponding positive evaluation content. This can be combined with the positive and positive incentive mechanisms mentioned above to generate positive evaluation content. For example, the positive evaluation content could be "This student's lecture was very good, and everyone was listening carefully."
[0201] The generated positive evaluation content may be in a variety of data modes, such as outputting encouraging sentences through the user interface of the target application, or simulating a target virtual image to emit a voice containing encouraging content.
[0202] In this application, neural network algorithms such as convolutional neural networks (CNN) or human posture analysis algorithms such as OpenPose can be used to analyze the facial expressions and body movement characteristics of each user. Recurrent neural networks (RNN) and natural language processing technologies can also be combined to obtain the user's audio features and semantic features to comprehensively determine the emotion category of each user and judge whether emotional resonance points have occurred. For example, by nodding, clapping, and other actions by multiple users, it is determined that there is currently emotional resonance.
[0203] In addition, based on each user's emotional feature data, knowledge graph data, and user data and questionnaire feedback, the parameters of each algorithm or model for determining the emotional model point can be dynamically adjusted to more accurately determine the emotional resonance point.
[0204] In another possible scenario, when multiple users are divided into multiple user groups, for any one of the multiple users, when it is determined that the user belongs to a set type of user based on the user's emotional characteristic data, the user's interaction data is anonymized, and / or session data is exchanged through a target communication channel between the user and the target virtual image.
[0205] The set user type can be a user with set emotional characteristics, that is, a user belonging to a set user type. For example, the set user type can be a user with low self-esteem, or a user with timidity, etc. The set user type can be one or more, without limitation.
[0206] There are many specific ways to determine whether a user belongs to a set type of user based on the user's emotional characteristic data. For example, the emotional characteristic data may have keywords associated with the set type of user, and then determine that the user belongs to the set type of user. For another example, the present application can pre-train a classification model for each set user type using the user's emotional data of the user type. Accordingly, for each user, the classification model can be used to determine whether the user belongs to the user type. If so, the user belongs to the set type of user corresponding to the user type.
[0207] When the user belongs to a set type of user, the user's interaction data is anonymized so that other users cannot see the owner information of the interaction data entered by the user (such as the user's name, nickname or student number, etc.). In this way, even if the interaction data entered by the user contains errors, it can avoid increasing the user's psychological burden.
[0208] The target communication channel can be an independent communication channel between the user and the target avatar, such as a private messaging channel or email. By exchanging conversation data between the user and the target avatar through the target communication channel, it is possible to send encouraging messages to the user or point out any problems the user may have in their learning process, thereby preventing the user from feeling embarrassed or awkward.
[0209] For example, taking the scenario of assisted teaching based on the target application as an example, after obtaining multi-modal data such as student voice, text, and facial images based on the target application, if the analysis shows that the student's voice has hesitant and depressed intonation, the text contains self-negative words, and the facial expressions have emotional characteristics such as avoidance eyes and frustration, it can be determined that the student has inferiority complex. Once the student's inferiority complex is detected, the interaction strategy of the target virtual image can be automatically adjusted to avoid mentioning the student's learning deficiencies. For example, during a classroom discussion, if a student is found to exhibit emotional characteristics related to inferiority, the virtual teacher will not directly point out their mistakes or deficiencies. Instead, the virtual teacher can provide personalized advice and encouragement through private messages to protect the student's emotional privacy.
[0210] For easier understanding, see Figure 9 , Figure 9 A schematic diagram showing an interactive interface of the target application in this application is shown.
[0211] exist Figure 9 In addition to the information display area 901, the information input bar 902, the displayed target virtual image 903, the auxiliary function operation area 904 and the information feedback area 905 corresponding to the target virtual image, the interactive interface also includes an icon 906 for the private message channel.
[0212] When the target avatar provides conversation data to the user via the target communication channel, the icon for the private messaging channel will indicate the presence of unread messages, such as displaying a number representing the number of unread messages. When the user clicks the icon for the private messaging channel, the conversation data provided by the target avatar to the user will be displayed.
[0213] In addition, as described in the previous example, the information feedback area 905 can present the target virtual image's evaluation or encouraging words for the user, and the content displayed in the information feedback area can be consistent with the content played by the target virtual image through voice.
[0214] In another possible scenario, when the target user includes multiple users, the present application can also update the output parameters of the target avatar based on the emotional characteristic data of the multiple users. For example, based on the emotional characteristic data of the multiple users, the comprehensive learning performance or emotional state of the multiple users can be determined, and the expression, movement, and output voice of the target avatar can be updated.
[0215] In another possible scenario, when the target user includes multiple users, the present application can also generate corresponding target virtual images for different user groups, or update the output parameters of the target virtual image based on the group emotional characteristic data of the user group.
[0216] The target virtual images corresponding to different user groups may be different.
[0217] The group emotion characteristic data is a comprehensive representation of the emotion characteristic data of each user in the user group.
[0218] For example, for each user group, the output parameters of the target virtual image can be updated based on the emotional feature data or knowledge graph data of all users in the user group.
[0219] In addition, in any of the above embodiments of the present application, the present application can also combine the user's (or target user's) suggestions and opinions on the target virtual image or target application and the user's feedback information on the response to the target application (for example, whether the user actively participates in asking questions in the teaching scenario, changes in the frequency of answering questions, etc.), etc., and can dynamically adjust the parameters of relevant algorithms and models used to determine emotional feature data, environmental feature data, output parameters of the target virtual image, and emotional resonance points, so as to more accurately generate and update the target virtual image and provide services to users more effectively and reasonably.
[0220] Of course, the present application can also combine the user's interaction data with the target application, emotional state, etc. to update the vocabulary related to the user and the target virtual image.
[0221] On the other hand, the present application also provides an electronic device. Figure 10 , which shows a schematic diagram of the composition structure of the electronic device, which includes at least one processor 1001 and at least one processing model that can run on the processor. The processing model can be called by the target application to perform at least one of the following:
[0222] In response to a target triggering event, obtaining target interaction data between a target user and the electronic device;
[0223] generating a target virtual image capable of interacting with the target user based on the target interaction data;
[0224] updating output parameters of the target virtual image based on the obtained target reference data, wherein the target reference data may or may not be from the target user, and the output parameters can reflect the interaction process between the target user and a target application of the electronic device, wherein the target application is an application capable of invoking at least one processing model to at least generate or update the target virtual image;
[0225] Among them, under different triggering events, the target interaction data obtained is different.
[0226] The electronic device may further include a memory 1002 for storing programs required for the processor to perform operations.
[0227] It is understandable that the electronic device may further include a display unit 1003 and an input unit 1004 .
[0228] Of course, the electronic device may also have Figure 10 There is no limitation to more or fewer components.
[0229] A computer program product is also provided in an embodiment of the present application, including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any data processing method provided in the embodiment of the present application.
[0230] A computer-readable storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any data processing method provided in the embodiment of the present application.
[0231] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.
[0232] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods described in each embodiment of the present application.
[0233] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0234] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a training device or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website, a computer, a training device or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
Claims
1. A data processing method, comprising: In response to a target triggering event, obtaining target interaction data between a target user and the electronic device; generating a target virtual image capable of interacting with the target user based on the target interaction data; updating output parameters of the target virtual image based on the obtained target reference data, wherein the target reference data may or may not be from the target user, and the output parameters can reflect the interaction process between the target user and a target application of the electronic device, wherein the target application is an application capable of invoking at least one processing model to at least generate or update the target virtual image; Among them, under different triggering events, the target interaction data obtained is different.
2. The data processing method according to claim 1, wherein: Obtaining target interaction data between a target user and an electronic device includes at least one of the following: Based on the type of the target trigger event, at least one of target input data input by the target user to the target application, application data of the target application, or obtained behavior feature data of the target user is used as the target interaction data; The target interaction data is acquired in a corresponding acquisition manner based on the type of the target triggering event.
3. The data processing method according to claim 1, wherein generating a target virtual image capable of interacting with the target user based on the target interaction data comprises: Identifying a domain category to which the target interaction data belongs and a user intent represented, wherein the domain category is determined based at least on an attribute of a target content in the target interaction data; At least one processing model is called based on the domain category and the user intention to generate a target virtual image for interacting with the target user.
4. The data processing method according to claim 1 or 2, wherein generating a target virtual image capable of interacting with the target user based on the target interaction data comprises: Invoking at least one processing model to generate and process the obtained behavior characteristic data of the target user to obtain a user virtual image that matches the behavior characteristic data; as well as, Based on the application data of the target application and / or the target input data inputted by the target user to the target application, a target character virtual image capable of interacting with the user virtual image in a target virtual scene is generated.
5. The data processing method according to claim 1 , wherein updating the output parameters of the target virtual image based on the obtained target reference data comprises: Identifying emotional characteristic data and / or environmental characteristic data of a target user during interaction with the target application; At least one of the constituent elements, character positioning, posture and movement, emotional expression, and sound parameters of the target virtual image is updated based on the emotional feature data and / or the environmental feature data.
6. The data processing method according to claim 5, wherein the identifying of the target user's emotional characteristic data and / or environmental characteristic data during the interaction with the target application comprises at least one of the following: Identifying emotional attributes and / or environmental attributes carried in target input data input by a target user to a target application and / or feedback data of a response result of the target application, and obtaining the emotional feature data and / or environmental feature data based on the emotional attributes and / or the environmental attributes; Obtaining behavioral characteristic data of the target user during the interaction process, and identifying the behavioral characteristic data to obtain the emotional characteristic data; Obtaining environmental change data of the spatial environment in which the electronic device is located, and identifying the environmental change data to obtain the environmental characteristic data; and / or, The updating of the output parameters of the target virtual image based on the obtained target reference data further includes: updating output parameters of the target virtual image based on attribute change information of the interaction content between the target user and the target application; The output parameters of the target virtual image are updated based on usage change information of the usage scenario or usage purpose of the target application.
7. The data processing method according to claim 1 or 5, wherein: In the case where the target user includes multiple users, the data processing method further includes: Grouping the multiple users based on the emotional feature data and / or knowledge graph data of the multiple users to obtain multiple user groups; Corresponding guidance data is generated based on the tag data corresponding to each user group, so as to guide the user group to complete the target task based on the corresponding guidance data.
8. The data processing method according to claim 7, further comprising at least one of the following: Regrouping based on changes in the emotional feature data and / or knowledge graph data of users in each user group; Generate corresponding positive evaluation content based on the emotional resonance points of multiple users in the user group.
9. The data processing method according to claim 7, further comprising at least one of the following: When it is determined based on the user's emotional characteristic data that the user belongs to a set type of user, anonymizing the user's interaction data and / or exchanging conversation data through a target communication channel between the user and a target avatar; updating output parameters of the target avatar based on the emotional characteristic data of the plurality of users; Generate corresponding target virtual images for different user groups, or update the output parameters of the target virtual images based on the group emotional characteristic data of the user groups.
10. An electronic device comprising at least one processor and at least one processing model capable of running on the processor, wherein the processing model can be called by a target application to perform at least one of the following: In response to a target triggering event, obtaining target interaction data between a target user and the electronic device; generating a target virtual image capable of interacting with the target user based on the target interaction data; updating output parameters of the target virtual image based on the obtained target reference data, wherein the target reference data may or may not be from the target user, and the output parameters can reflect the interaction process between the target user and a target application of the electronic device, wherein the target application is an application capable of invoking at least one processing model to at least generate or update the target virtual image; in, Under different triggering events, the target interaction data obtained is different.