Information processing system

CN122795902APending Publication Date: 2026-09-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610248360.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-19
Filing Date
2026-03-02
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0003]在现有技术中,虽然已经存在用于与用户进行对话的虚拟角色或头像系统,但仍存在以下主要问题:其一,传统头像多以固定预设的外观和性格为主,系统无法根据用户输入的自然语言提示动态地解析用户需求,从而生成与用户个性、喜好相匹配的头像外观与性格配置,难以提供高度个性化的互动体验

Benefits of technology

1. 服务器在“用户配置数据库”中存储每个用户的属性向量,属性向量包括用户的基本属性字段(例如年龄段、性别类别的编码)、偏好标签向量(如对“温柔”“活泼”性格的偏好权重)以及历史情感分布统计(如最近一段时间内的“喜悦”“悲伤”“压力”等比例)。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122795902A_ABST
    Figure CN122795902A_ABST
Patent Text Reader

Abstract

The application provides an information processing system. An information processing system, characterized by comprising: a processor; wherein the processor is configured to: parse a prompt sentence received from a user, generate feature data for indicating determination of an avatar appearance and personality; send the generated feature data to a terminal of the user to visually display the avatar on the terminal of the user; and infer an emotional state of the user, and parse a prompt for generating an avatar reaction corresponding to the emotional state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology disclosed herein relates to an information processing system. Background Technology

[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot response to the user's speech.

[0003] While existing technologies exist for virtual characters or avatars to interact with users, they still suffer from several major problems: First, traditional avatars are mostly based on fixed, preset appearances and personalities. The system cannot dynamically analyze user needs based on natural language input to generate an avatar appearance and personality that matches the user's personality and preferences, making it difficult to provide a highly personalized interactive experience. Second, existing systems generally lack effective mechanisms for recognizing and responding to users' emotional states. They cannot intelligently analyze user input to infer the user's current emotional state and generate an avatar response that matches that emotion, resulting in a lack of emotional resonance and psychological comfort in the interaction between the avatar and the user. Third, due to the lack of continuous perception of user emotional changes and targeted interaction strategies, existing technologies struggle to effectively alleviate users' loneliness over long-term use, especially for emotionally sensitive users or those who spend long periods alone, whose psychological support and companionship needs cannot be fully met. Therefore, there is a need for a system that can analyze user prompts based on natural language, automatically generate avatar feature data, and generate corresponding avatar responses based on the user's emotional state, so as to achieve human-computer interaction with greater emotional understanding and companionship, thereby effectively alleviating users' loneliness. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention provides an information processing system comprising a processor, wherein: the processor is configured to parse prompts received from a user and generate feature data using prompts indicating the appearance and personality of an avatar; the processor is further configured to send the generated feature data to a user's terminal for visual display of the avatar on the user's terminal, enabling the avatar's appearance and personality to be dynamically customized based on the user's preferences and needs expressed through natural language. Furthermore, the processor is also configured to infer the user's emotional state and parse prompts for generating avatar responses corresponding to the emotional state, thereby generating appropriate avatar responses based on the user's current emotion and achieving emotionally adaptive interaction.

[0005] Preferably, the processor is configured to parse the prompt statement using natural language processing technology to understand the prompt indicating an change in avatar appearance and / or personality. This allows users to adjust the avatar's appearance and personality traits solely through natural language input, without requiring complex manual settings, thus improving the system's usability and personalization. Furthermore, the processor is configured to generate corresponding avatar responses based on the user's emotional state and alleviate loneliness through interaction between the user and the avatar. Specifically, when the processor detects the user is in a state of sadness, anxiety, or loneliness, it generates an avatar response that offers comfort, encouragement, or companionship; when the processor detects the user is in a positive emotional state such as joy or excitement, it generates an avatar response that shares the user's joy. Through these configurations, the present invention can achieve intelligent recognition and emotional response to the user's emotional state while dynamically customizing the avatar's appearance and personality, thereby constructing a human-computer interaction system with emotional companionship capabilities, effectively alleviating the user's loneliness and enhancing the overall interactive experience.

[0006] "System" refers to a collection of overall technical solutions, including at least one processor and optional memory, communication interfaces and / or user terminals, for performing functions such as prompt statement parsing, feature data generation, data transmission and avatar response generation.

[0007] "Processor" means hardware or a combination thereof capable of executing program instructions to perform data processing and logical operations, including but not limited to a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a programmable logic device (FPGA), or any combination thereof.

[0008] "Prompt statements" refer to text information that is input by the user in natural language and received by the system, or text information that is converted from voice input by speech recognition. This text contains semantic content used to indicate the appearance of the avatar, personality and / or the generation of reactions, etc.

[0009] "Prompt" refers to semantic instructions or parameter information obtained from the prompt statement through parsing, which are used to instruct the system to perform specific operations (including but not limited to determining or changing the appearance of the avatar, the personality of the avatar, and generating the avatar reaction).

[0010] “Feature data” refers to a set of structured data or parameters generated based on the parsed prompts, used to characterize the appearance and / or personality traits of an avatar. This data can be used by user terminals or servers to drive the visual display and behavioral performance of the avatar.

[0011] "User terminal" refers to a device operated by the user and communicating with the server to display avatars, receive user input, and present system feedback, including but not limited to smartphones, tablets, personal computers, smart TVs, wearable devices, or other electronic devices with display and interactive functions.

[0012] A "profile picture" refers to a virtual image that is visualized on a user's terminal by the system based on feature data. This virtual image has perceptible appearance features and / or personality features and is able to interact with the user.

[0013] "Appearance" refers to the visual characteristics of an avatar, including but not limited to body type (such as animal or human form), color, facial features, clothing style, and other visual elements.

[0014] "Personality" refers to the behavioral style and expressive characteristics set by the system for an avatar during the interaction process, including but not limited to talkativeness, emotional expression, tone of voice, positivity, and politeness.

[0015] "Emotional state" refers to the system's prediction of a user's psychological or emotional state at a specific point in time based on the analysis of user input and / or related data, including but not limited to states such as neutral, happy, sad, anxious, angry, or lonely.

[0016] "Avatar response" refers to the response generated by the system after inferring the user's emotional state, which is manifested on the user's terminal by the avatar, including but not limited to text replies, voice output, facial expression changes, action changes, or any combination of the above.

[0017] Natural Language Processing (NLP) technology refers to a set of computer technologies used to automatically analyze and understand user prompts, including but not limited to word segmentation, part-of-speech tagging, syntactic analysis, semantic parsing, sentiment analysis, and dialogue intent recognition, in order to extract semantic information related to avatar control and response generation from natural language. Attached Figure Description

[0018] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.

[0019] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.

[0020] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.

[0021] Figure 4This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.

[0022] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.

[0023] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and head-mounted terminal according to the third embodiment.

[0024] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.

[0025] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.

[0026] Figure 9 This represents an emotion map that maps multiple emotions.

[0027] Figure 10 This represents an emotion map that maps multiple emotions.

[0028] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.

[0029] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.

[0030] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.

[0031] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation

[0032] Hereinafter, an example of an implementation of the system to which the technology of this disclosure relates will be described with reference to the accompanying drawings.

[0033] First, let me explain the terminology used in the following instructions.

[0034] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.

[0035] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.

[0036] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.

[0037] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.

[0038] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.

[0039] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.

[0040] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.

[0041] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0042] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.

[0043] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.

[0044] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0045] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.

[0046] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.

[0047] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0048] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0049] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.

[0050] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.

[0051] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."

[0052] With the development of generative artificial intelligence models and virtual avatar display technology, users can interact with virtual objects through information terminals. However, existing technologies suffer from the following problems: First, servers typically generate the appearance or simple personality tags of virtual objects based solely on a single static input, lacking a mechanism for deep semantic analysis of users' natural language prompts. This makes it difficult for virtual objects to accurately reflect users' multidimensional preferences and nuanced personality needs, thus reducing the degree of customization. Second, servers often treat generative artificial intelligence models as "black boxes," directly using user text as input. They lack structured construction and dynamic adjustment methods for prompts, resulting in poor consistency, controllability, and interpretability of the generated results, making it difficult to maintain an expected interaction style over the long term. Third, existing systems often generate responses based solely on the current sentence during the dialogue phase, lacking modeling of changes in the user's emotional state over time and a technical solution for combining these emotional changes with the virtual object's predetermined personality attributes for response control. This limits the effectiveness of virtual objects in scenarios such as companionship, comfort, and alleviating feelings of isolation. Fourth, during multiple rounds of interaction, the server typically does not dynamically update the feature data of virtual objects based on continuous input, making it impossible to continuously iterate and customize the appearance and personality attributes of virtual objects. As a result, users experience insufficient immersion and intimacy when using the software for an extended period.

[0053] Therefore, an improved computer implementation scheme is needed. On the server side, this scheme involves structured parsing of the user's natural language input, constructing various prompts for generative artificial intelligence models, and combining modeling of the user's emotional state and interaction history. This would improve the quality of virtual object generation and interactive experience, while also improving the server's calling method for generative artificial intelligence models and data processing flow. This would enable high-precision customization and continuous optimization of virtual objects, thereby substantially improving the functional performance of the human-computer interaction system through technical means.

[0054] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.

[0055] In this invention, the server includes: a processing unit configured to acquire a prompt statement containing natural language input information from a user-operated information terminal, parse the prompt statement, and extract preference information and demand conditions from it; a processing unit configured to construct a generation prompt statement for input to a generative artificial intelligence model based on the preference information and demand conditions, and call the generative artificial intelligence model to generate feature data representing the appearance attributes, personality attributes, and behavioral attributes of a virtual object; a processing unit configured to send visual representation information from the feature data to the information terminal so as to visually display the virtual object on the information terminal; and a processing unit configured to... The information terminal continuously acquires user dialogue input and physiological information or operation history to infer the user's emotional state, and constructs response generation prompts reflecting the emotional state and the personality attributes of the virtual object. A processing unit calls the generative artificial intelligence model to generate reaction data representing the virtual object's response content, facial expressions, or actions. A processing unit is configured to send the reaction data to the information terminal, causing the information terminal to present it as a virtual object speaking and displaying changes in facial expressions or actions. Furthermore, it constructs regenerated prompts based on the user's additional prompts and uses the generative artificial intelligence model to update the feature data to continuously customize the virtual object. This allows for structured analysis of natural language input and multi-level prompt construction on the server side, improving the controllability and consistency of the generative artificial intelligence model's output. Through dynamic modeling of the user's emotional state and interaction history, a closed loop is formed between the virtual object's appearance and personality generation, dialogue response, and subsequent customization. This improves the generation quality and interaction effect of the virtual interaction system at the computer technology level, effectively enhancing human-computer interaction performance and alleviating user isolation.

[0056] "Information processing device" refers to an electronic device that includes at least one processor, memory, and communication interface, used to execute program instructions, process input data, and interact with information terminals through a communication network.

[0057] The "processing unit" refers to the software functional module or logical unit executed by the processor, which is used to perform processing operations such as parsing input data, extracting features, constructing prompt statements, and calling generative artificial intelligence models according to a predetermined program.

[0058] "Information terminal" refers to an electronic device operated by a user for inputting natural language information, displaying virtual objects, and presenting interactive content, including but not limited to smartphones, tablet computers, personal computers, and other terminal devices with display functions.

[0059] "User" refers to the subject that interacts with the system through an information terminal, provides natural language input, preference information and operation instructions, and receives virtual object display and response content, usually a human individual.

[0060] "Natural language input information" refers to text information expressed in the form of human natural language or information converted into text through speech recognition, used to describe the user's preferences, needs, emotions, or requirements for virtual objects.

[0061] "Prompt statements" refer to a set of natural language text or structured instructions constructed to drive generative artificial intelligence models to perform specific generative tasks, which include generation goals, constraints, and semantic guidance information.

[0062] "Preference information" refers to semantic information extracted from user input that reflects the user's preferences in appearance style, personality type, interaction method, etc., and is used to guide the customization of virtual object features.

[0063] "Requirement conditions" refer to the requirements or restrictions imposed on the functionality, presentation, or interaction mode of virtual objects, extracted from user input or settings, including but not limited to style constraints, interaction frequency, and content preferences.

[0064] "Generation prompt statements" refer to prompt statements constructed by the processing department based on preference information and demand conditions after parsing user input. These prompt statements are used as input to the generative artificial intelligence model to instruct the generation of data related to the appearance, personality, and behavioral attributes of virtual objects.

[0065] "Generative artificial intelligence models" refer to artificial intelligence models trained using machine learning techniques that can automatically generate text, images, or other data content based on input prompts, including but not limited to language models and image generation models based on deep learning.

[0066] "Virtual objects" refer to digital characters or digital images that are presented on information terminals in the form of images, animations or other forms and interact with users under system control.

[0067] "Appearance attributes" refer to the characteristic information used to describe the visual appearance of virtual objects, including but not limited to gender orientation, hairstyle, clothing, facial expression style, and overall visual style.

[0068] "Personality attributes" refer to information used to describe the personality traits exhibited by virtual objects in dialogue and behavior, including but not limited to personality dimensions such as extroversion or introversion, quietness or liveliness, gentleness or rationality.

[0069] "Behavioral attributes" refer to the characteristic information used to describe the action patterns and behavioral tendencies of virtual objects during the interaction process, including but not limited to speaking frequency, proactive greeting methods, rhythm of facial expression changes, and types of posture and movement.

[0070] "Feature data" refers to a structured data set output by a generative artificial intelligence model, used to define and control the appearance, personality, and behavioral attributes of virtual objects. It may include text descriptions, numerical vectors, and resource index information used for rendering.

[0071] "Visual representation information" refers to image data, image generation parameters, or rendering-related descriptive information extracted from feature data and used to generate the visual presentation of virtual objects on information terminals.

[0072] "Dialogue input" refers to natural language text information or text information converted from speech recognition provided by the user through the information terminal during the interaction with virtual objects, which drives the system to generate response content.

[0073] "Physiological information" refers to biological or physiological data acquired through sensors or external interfaces to infer a user's emotional state, including but not limited to heart rate, skin conductance response, and facial expression analysis results.

[0074] "Operation history" refers to the sequence of user interaction behaviors recorded by the system on the information terminal, including but not limited to click records, input records, dwell time, and session frequency, which are used to analyze user status and behavior patterns.

[0075] "Emotional state" refers to the user's current psychological or emotional state inferred from dialogue input, physiological information, and operation history, including but not limited to emotional labels or continuous emotional values ​​such as happiness, sadness, anxiety, and loneliness.

[0076] "Response generation prompts" refer to prompts constructed by the processing department after estimating the user's emotional state and considering the personality attributes of the virtual object, used to instruct the generative artificial intelligence model to generate dialogue responses and corresponding expressions or actions.

[0077] "Response data" refers to structured data generated by generative artificial intelligence models based on response prompts, used to control the speech content, facial expressions, and actions of virtual objects.

[0078] "Speech display" refers to the display format on the information terminal's display interface that presents the virtual object's response content in the form of text, voice, or a combination of both.

[0079] "Regeneration prompt statement" refers to a prompt statement constructed by the processing department based on additional user input or new preference information, after the virtual object has been generated and interacted with by the user, to update the feature data of the virtual object.

[0080] "Continuous customization" refers to the process of dynamically updating and iterating the appearance, personality, and behavioral attributes of virtual objects by repeatedly using regenerable prompts and generative artificial intelligence models, in order to gradually improve the matching degree between virtual objects and user preferences.

[0081] This invention will describe the software and hardware structure for collaborative operation among the server, terminal, and user. It will also illustrate the data structure, algorithm flow, and the resulting improvements in computer technology by combining the specific construction methods of generative artificial intelligence models and prompt statements. The following embodiments are merely illustrative and do not limit the technical scope of this invention.

[0082] I. Overall System Composition A server may include at least one processor, main memory, non-volatile storage, a network interface, and an optional graphics processing unit. The processor may be a general-purpose central processing unit, the main memory may be random access memory, and the non-volatile storage may be a solid-state drive or a hard disk drive. The network interface may be an Ethernet interface or a wireless communication module. The server may run backend applications and generative artificial intelligence model services on an operating system (such as a Linux-based server operating system).

[0083] A terminal may include an electronic device with a display, input unit, communication module, and local storage, such as a smartphone, tablet, or personal computer. The terminal can run client applications or web browsers on a mobile or desktop operating system to collect user input and display virtual objects.

[0084] Users provide natural language input information to the server through the terminal and receive virtual object displays and interactive content returned by the server.

[0085] II. Server-side software modules and data structures The server can run application server programs on top of an operating system. These application server programs can be implemented using application frameworks, such as scripting language-based backend frameworks (like a general-purpose web framework) or runtime environment-based backend frameworks. The server may further include the following software modules: 1. Input parsing module The server can be configured with an input parsing module to receive natural language input from the terminal and convert it into an intermediate data structure. The input parsing module may include: (1) Text preprocessing submodule The server can perform string cleaning, unified encoding conversion, and punctuation standardization on user natural language input, converting the input into standardized text. The server can store the text in a key-value store with "session identifier" as the key for subsequent processing.

[0086] (2) Word segmentation and tagging submodule The server can use natural language processing tools (such as word segmentation tools based on dictionaries and statistical models, or sequence labeling models based on bidirectional recurrent neural networks) to segment and tag the text. The server can represent the result as a sequence structure, where each element includes a word, a part-of-speech tag, and its position index in the original text.

[0087] (3) Preference and Demand Extraction Submodule The server can be configured with template rule sets and lightweight classification models to map segmented sequences to "preference information" and "demand conditions". For example, the server can identify "cat" as the "preference object" in "I like cats, quiet personality, preferably a little gentle", and "quiet" and "gentle" as "personality tags".

[0088] The server can assign a discrete identifier to each tag and further generate a numerical attribute vector, such as a vector of length N, where each dimension corresponds to a predefined feature (such as animal type, extraversion, gentleness, initiative, etc.), which is written to the corresponding position through table lookup or weighted mapping.

[0089] 2. Prompt Statement Construction Module The server can be configured with a prompt statement construction module to generate various types of prompt statements based on extracted preference information, demand conditions, and historical interaction information, for input into the generative artificial intelligence model. This module may include: (1) Generate the part constructed using prompt statements The server can fill the extracted semantic information into a predefined prompt template. For example, the server can generate the following text as a prompt statement: "Generate a virtual avatar based on the following user preferences: The user likes cats, has a quiet and gentle personality. Please describe the appearance of the virtual avatar (including approximate gender, age, and clothing style), as well as its personality traits and how it interacts with the user, in simplified Chinese." For image generation, the server can construct prompts for image generation, such as: "A virtual character with a cat-like temperament, gentle and quiet appearance, simple clothing, soft color scheme, anime style, upper body image, and simple background." The server encodes discrete labels and numerical attribute vectors into the aforementioned prompt text or sends them as additional fields to the multimodal model, enabling generative artificial intelligence models to utilize this structured information in a continuous space.

[0090] (2) Response generation using prompt statements construction part The server can generate response prompts during the dialogue phase by combining "current user message," "historical dialogue summary," "emotional state vector," and "virtual object personality attributes." For example, the server can construct: "You are a virtual avatar with a quiet and gentle personality, possessing a cat-like quality, and a good listener. Your task is to accompany and comfort the user, speaking softly and not too loudly. Please answer the user's next question in Simplified Chinese." And append the conversation history and current user message after it, for example: User: I'm in a bad mood today, can you talk to me? The server can embed numerical values ​​such as "sadness" and "anxiety" from the emotion state vector into prompt statements through descriptive statements or specific control markers, thereby achieving fine-grained control over the output style of generative artificial intelligence models.

[0091] (3) Regenerate the part constructed using prompt statements The server can reconstruct the prompt statement when the user requests further customization. For example, if the user enters, "Make this avatar quieter," the server can generate: "Based on the existing virtual avatar settings, adjust its personality to make it quieter and less likely to interrupt the user. Please redescribe its personality traits and interaction methods while maintaining the original appearance style." 3. Generative Artificial Intelligence Model Module The server may include at least one generative artificial intelligence model service. This service may be deployed on a server node equipped with a graphics processing unit. The model may include: (1) Text generation model The server can use a neural network based on a self-attention mechanism, such as a transformer model with a multi-layered encoder-decoder structure. The model can include several encoding layers and several decoding layers, with the number of attention heads and the dimension of hidden layers being predetermined values. The server can encode prompts into a sequence of embedding vectors, pass them through multi-head attention and feedforward network layers, and generate virtual object descriptions or dialogue response text.

[0092] The model can use a loss function (such as cross-entropy loss) during the training phase to update the network weights through backpropagation. The server can employ data augmentation methods during training, such as paraphrasing user preference descriptions or randomly masking parts of words, to improve the model's robustness to diverse inputs.

[0093] (2) Image generation model The server can use a diffusion-based image generation network. This model can include a text encoder, a noise prediction network, and a decoding module. The server inputs a prompt statement into the text encoder to obtain a high-dimensional embedding vector; then, starting with Gaussian noise in the latent space, it generates latent vectors through a multi-step denoising process; finally, the decoding module generates the corresponding image. During training, mean squared error loss or perceptual loss functions can be used, and data augmentation (such as random cropping and color perturbation) can be employed to improve generalization ability.

[0094] (3) Personality and Emotional Control Mechanisms The server can add control tag vectors to the model input to map the personality attributes of the virtual object and the current emotional state of the user to additional control channels. This can be done by concatenating a control vector of length K into the model input embedding, or by adding special control tags to the prompt statements.

[0095] With this design, the server can adjust the style and emotional tone of the output under the same model structure, achieving fine-grained control, without having to train a separate model for each personality, thereby reducing computational resource consumption and improving response speed.

[0096] 4. Status Management and Data Storage Module The server can be configured with a database or key-value store system to record user session states, emotional state time series, and virtual object attribute versions. (1) User session records The server can assign a session identifier to each user and use it as the key to store a summary of the conversation history. To control storage and computational overhead, the server can use a windowing mechanism or a summarizing algorithm (such as an attention-based summarizing network) to compress long conversations, retaining only the important statements used for contextual understanding.

[0097] (2) Emotional state vector storage The server can encode the presumed emotional state after each conversation into a multi-dimensional vector (e.g., dimensions representing happiness, sadness, anger, anxiety, loneliness, etc.), and store it in a time-series database with a timestamp. When generating a response, the server can use the emotional vectors from recent moments to calculate a weighted average or trend indicator as a dynamic adjustment parameter.

[0098] (3) Virtual object configuration version The server can assign a version number to the feature data of each updated virtual object and store the appearance, personality, and behavioral attributes of each version in the database. Based on user selection or system policy, the server can restore or switch to an existing version, enabling rollback and multi-style management.

[0099] III. Terminal-side software structure and display control The terminal can run client applications. Client applications may include: 1. Input Interface Module The terminal can provide a text input field and an optional voice input button. The terminal can locally invoke a speech recognition service to convert the speech signal into text before sending it to the server. The terminal can perform length checks and basic filtering on the input content, but does not perform complex semantic analysis to reduce the terminal's workload.

[0100] 2. Communication Module The terminal can establish a secure connection with the server through the Transmission Control Protocol / Internet Protocol stack (e.g., using Transport Layer Security). The terminal can send user input, session identifiers, and local device information in the form of structured requests, and receive characteristic data and response data returned by the server.

[0101] 3. Virtual Object Display Module The terminal can parse the appearance description and image resource address returned by the server. For two-dimensional images, the terminal can draw the avatar image on the display using an image rendering engine. For three-dimensional images, the terminal can load models and animation resources in a 3D graphics engine and drive skeletal animation or facial expression blending shapes according to facial expression or motion instructions provided by the server.

[0102] 4. Dialogue Display and Motion Control Module The terminal can display the server's text reply in the form of bubbles in the chat interface, and can optionally call a text-to-speech engine to convert the text into speech for playback. The terminal can trigger corresponding animation sequences locally based on the "action tags" (such as "soft_smile" and "wave_hand") in the response data to achieve synchronized presentation of facial expressions and actions.

[0103] IV. Examples of User Interaction and Prompt Statements Users can input various natural language prompts on the terminal for initial generation or subsequent customization of virtual objects. Examples include: "I like cats, with quiet personalities, preferably a bit gentle." "I like dogs. They should be outgoing, talkative, and have a cute profile picture." "I've been feeling a bit lonely lately, and I'd like a profile picture that's supposed to listen, but not too noisy." "Let this profile picture be quieter." The server can parse these prompt statements separately and construct corresponding prompt statements for generating responses, prompt statements for generating responses, or prompt statements for regenerating responses, as shown in the examples in the aforementioned implementation.

[0104] V. Explanation of Technical Effects and Causal Relationship Through the aforementioned structured processing and multi-layered prompts, the server can bring about several improvements at the computer technology level compared to simply inputting raw user text directly into a generative artificial intelligence model: 1. Improved accuracy and consistency The server maps natural language input to a low- or medium-dimensional feature space using intermediate representations such as "preference information," "demand conditions," "personality attribute vectors," and "emotional state vectors," then guides the generative AI model to generate results. This feature-based processing reduces the model's sensitivity to noisy input, making the appearance and personality of virtual objects more stable and maintaining a consistent style throughout multi-turn dialogues.

[0105] 2. Improved computational efficiency and response speed The server compresses emotional state time series data and dialogue history summaries to avoid transmitting and processing complete historical data with each generation, thereby reducing the model input length and computational cost. Since the complexity of generative AI models is approximately quadratically or linearly related to the input length, this truncation and summarization strategy can significantly reduce computation time and improve response speed.

[0106] 3. Reduced communication load and storage pressure The server reduces network communication load by compressing the settings, personality, and behavioral attributes of virtual objects into feature data and concise descriptions, rather than transmitting large amounts of resources in every interaction. The terminal only downloads or updates image and model resources when necessary, thus reducing bandwidth consumption. Furthermore, through session summaries and version management, the server can store only critical states, significantly reducing storage requirements.

[0107] 4. Refined emotional response and reduced error By introducing multidimensional emotion state vectors and dynamic emotion regulation parameters, the server enables the generative AI model to adjust sentence structure and content based on quantitative values ​​such as "sadness level" and "loneliness level" during output, avoiding responses that do not match the user's state. Experiments demonstrate that this control mechanism reduces emotion response error by measuring the system's matching rate of emotion tags or user satisfaction scores.

[0108] 5. Rules and methods of non-human approaches The server constructs prompts using a combination of structured templates and vector injection. Instead of simply mimicking human "free description," it builds input based on the machine's optimal context organization strategy, enabling the model to utilize the context more effectively. This unconventional input method, centered on feature vectors and control tags, is a "machine-oriented" rule set designed for the internal computational characteristics of generative AI models, improving generation quality and stability with the same hardware resources.

[0109] VI. Model Training and Learning Methods Servers can train or fine-tune generative artificial intelligence models in an offline environment: 1. Training Data Composition The server can build a dataset containing "user preference description - virtual object setting" and "dialogue context - emotional state - response content", and encode information such as preferences, personality, and emotions into supervision signals during the annotation process.

[0110] 2. Learning Objectives and Loss Function For the text generation part, the server can use cross-entropy loss to make the model output distribution approximate the true target distribution. For the emotion control part, the server can add an auxiliary loss term to constrain the predicted label of the output text on the emotion classifier to be close to the target emotion label. For the image generation part, the server can use the noise prediction loss and perceptual consistency loss inherent in the diffusion model.

[0111] 3. Weight Update and Optimization Algorithm The server can use common gradient descent variant optimization algorithms to update model parameters and use techniques such as learning rate scheduling and gradient clipping to stabilize training.

[0112] 4. Decoding strategies in the reasoning phase The server can employ strategies such as temperature sampling, beam search, or kernel sampling during the inference phase to control the balance between output diversity and stability. By introducing "personality attribute vectors" and "emotional state vectors" into the decoding process, the server can change the bias direction of the probability distribution during decoding to ensure that the generated results conform to the set personality and emotion.

[0113] VII. Alternative Implementation Forms and Expansion Servers can employ various alternative architectures in different implementation forms: 1. Model architecture replacement The server can use a single large model to handle avatar generation and dialogue generation simultaneously, or deploy text generation models and image generation models separately. The server can also use pluggable control modules to encapsulate personality and emotion control logic into independent sub-networks.

[0114] 2. Terminal Local Inference Extension On terminals with strong computing power, the server can distribute some lightweight models to the terminal, allowing the terminal to perform simple character adjustments or basic dialogue generation locally, thus reducing the server load. In this case, the server can mainly handle complex generation and long-term state management.

[0115] 3. Multimodal input extension The server can be expanded to process multimodal inputs such as images, facial expressions, and videos simultaneously, enabling more accurate estimation of the user's emotional state. Correspondingly, the construction of the emotional state vector and cue statements can incorporate descriptions of visual cues, further improving response accuracy.

[0116] Through the aforementioned implementations, a complete data and control flow is formed among the server, terminal, and user. The system of this invention does not merely automate human work; rather, it substantially improves computer technology itself in terms of computational efficiency, generation accuracy, interaction stability, and resource management through structured processing of natural language input, machine-guided construction of prompts, controllable invocation of generative artificial intelligence models, and quantitative modeling of emotional states.

[0117] use Figure 11 The processing flow is explained.

[0118] Step 1: Users input natural language information through the terminal. Users can open the application interface on the terminal, enter a natural language description in the text input box, or speak their needs through the voice input button, such as: "I like cats, with a quiet personality, preferably a bit gentle." Input: The user's natural language text (or text converted from speech recognition).

[0119] The terminal performs preliminary processing on the string, including removing leading and trailing spaces, standardizing the encoding format, checking the length limit, and encapsulating the text together with the user identifier and session identifier into request data.

[0120] Output: Structured request data containing fields [user_id, session ID, natural language text].

[0121] Step 2: The terminal sends request data to the server. The terminal invokes the communication module to send the structured request data obtained in step 1 to the pre-defined interface of the server via a secure connection.

[0122] Input: Structured request data containing the user's natural language text.

[0123] The terminal uses network protocols to serialize (e.g., convert to JSON string), encrypt and encapsulate the data, and sends it to the server address via HTTP request.

[0124] Output: The request message transmitted to the server via the network.

[0125] Step 3: The server receives and parses the request message. When the server receives a request message from the terminal at the network interface, it forwards it to the backend application for processing.

[0126] Input: Request message from the terminal (including user text, user ID, session ID, etc.).

[0127] The server decodes and deserializes the message, parses the JSON string into an internal data object, checks whether the required fields exist, and extracts the natural language text and stores it in a temporary variable or session state structure.

[0128] Output: An internal data object containing the user's natural language text, user ID, and session ID.

[0129] Step 4: The server preprocesses the natural language text. The server performs text normalization and word segmentation on the natural language text obtained in step 3.

[0130] Input: User's natural language text.

[0131] The server calls the text preprocessing module to process the text: removing redundant spaces and control characters, converting full-width characters to half-width characters, and standardizing punctuation; then it calls the word segmentation tool to segment the text and perform part-of-speech tagging, converting the string into a sequence of words and their grammatical tags.

[0132] Output: Preprocessed data including lexical sequences, part-of-speech tags, and text position indices.

[0133] Step 5: The server extracts preference information and requirements. Based on the preprocessing results from step 4, the server identifies the user's preferred objects, personality traits, and needs.

[0134] Input: Preprocessed lexical sequence and part-of-speech tags.

[0135] The server uses a predefined set of rules and a lightweight classification algorithm to perform pattern matching and semantic classification on words and phrases, labeling words such as "cat," "quiet," and "gentle" as "preference objects" and "personality tags," and mapping these tags to internal codes. At the same time, it generates multi-dimensional attribute vectors (such as animal type dimension, extraversion dimension, gentleness dimension, etc.).

[0136] Output: Structured preference information, demand conditions, and corresponding attribute vectors.

[0137] Step 6: The server generates prompt statements. Based on the preference information and attribute vectors obtained in step 5, the server generates prompts to drive the generative artificial intelligence model.

[0138] Inputs: Preference information, requirement conditions, attribute vectors.

[0139] The server fills semantic tags into a predefined text template, concatenates the strings, and forms a prompt in natural language, such as: "Generate a virtual avatar based on the following user preferences: The user likes cats, has a quiet and gentle personality. Please describe the appearance of the virtual object (including approximate gender, age, and clothing style), as well as its personality traits and how it interacts with the user in simplified Chinese." Simultaneously, it can construct prompts for image generation, such as: "A virtual character with a cat-like temperament, a gentle and quiet appearance, simple clothing, soft color tones, anime style, upper body image, and a simple background." The server organizes these texts together with attribute vectors into a model input structure.

[0140] Output: Generative prompts for generative artificial intelligence models and their associated attribute vectors.

[0141] Step 7: The server calls a text generation model to generate virtual objects. The server uses a text sub-model of a generative artificial intelligence model to reason about the prompts in step 6, generating descriptions of the virtual object's appearance and personality.

[0142] Input: Generate a prompt statement and attribute vector.

[0143] The server segments and embeds the prompt statement, converting each word into a high-dimensional vector, and concatenates or weights the attribute vectors into the input vector. Then, it performs multi-layer self-attention calculation and feedforward operation in the transformer network, progressively calculating the output probability distribution, and generating descriptive text according to the set sampling strategy (such as temperature sampling or beam search), for example: "The virtual avatar is a young woman with a cat-like temperament, a refreshing appearance, smooth short hair, and wearing a simple sweater. She is quiet and gentle, likes to listen to the user's feelings, will not interrupt the user, and speaks in a soft tone." Output: Natural language text describing the appearance and personality attributes of the virtual object.

[0144] Step 8: The server calls an image generation model to generate a virtual object image (optional). Based on the prompt statement for image generation in step 6, the server calls the image sub-model of the generative artificial intelligence model to generate the avatar image.

[0145] Input: Image generation prompts and optional attribute vectors.

[0146] The server first maps the prompt statement into a text embedding vector using a text encoder. Then, it iteratively denoises the random noise in a diffusion model. At each step, a noise prediction network is used to correct the current image representation, gradually converging to a latent image representation that meets the text conditions. Finally, a decoding network restores the latent representation to a pixel image. The server saves the generated image file to a storage device and generates an accessible resource address.

[0147] Output: Image data of the corresponding virtual object and its access address.

[0148] Step 9: Server-generated feature data The server organizes the results of steps 7 and 8 into a unified feature data structure.

[0149] Input: Text description of the virtual object, image data address, attribute vector.

[0150] The server creates a feature data object, writes fields such as appearance description, personality description, behavioral parameters (such as speaking frequency and proactive greeting methods), and image URL into the object, and assigns a unique identifier to the feature data for subsequent session reference.

[0151] Output: Virtual object feature data containing appearance attributes, personality attributes, behavioral attributes, and resource reference information.

[0152] Step 10: The server sends the feature data to the terminal. The server returns the feature data generated in step 9 to the terminal that initiated the request.

[0153] Input: Virtual object feature data.

[0154] The server serializes the feature data, encapsulates it into a response message, and sends it to the terminal via the network interface; at the same time, it stores a copy of the feature data and its version number in a local database for subsequent updates.

[0155] Output: A response message (containing virtual object characteristic data) transmitted to the terminal.

[0156] Step 11: The terminal parses feature data and displays virtual objects. The terminal receives the response message returned by the server, parses the content, and presents virtual objects on the interface.

[0157] Input: A response message containing virtual object characteristic data.

[0158] The terminal deserializes the message, extracts the appearance description text and image URL; the terminal calls the image loading library to download the avatar image from the server or resource storage, and draws the image in a designated area of ​​the display screen; the terminal displays the appearance and personality description in text form in the description area, providing the user with an initial visual and textual impression.

[0159] Output: The virtual object image and text description displayed on the terminal interface.

[0160] Step 12: Users interact with virtual objects through the terminal. Users can enter their conversation in the chat interface on the terminal, such as: "I'm in a bad mood today, can you chat with me?", or express themselves via voice.

[0161] Input: New natural language dialogue text.

[0162] The terminal performs preprocessing on the text similar to step 1, including the current virtual object identifier (such as avatar_id) and a summary of the dialogue history, to generate dialogue request data.

[0163] Output: Structured data containing user dialogue text, virtual object identifiers, and a summary of the dialogue history.

[0164] Step 13: The terminal sends a dialogue request to the server. The terminal sends the dialogue request generated in step 12 to the server through the communication module.

[0165] Input: Structured data for the dialogue request.

[0166] The terminal serializes structured data into network packets and sends them to the server's designated dialog interface via a secure connection.

[0167] Output: The dialog request message that arrives at the server.

[0168] Step 14: The server estimates the user's emotional state. The server infers the user's current emotional state based on the latest dialogue input and historical interaction data.

[0169] Input: User's current dialogue text, conversation history summary, possible physiological information and operation history.

[0170] The server calls a sentiment analysis model to classify the text by emotion, mapping expressions such as "feeling bad" to the categories of sadness or frustration. At the same time, it performs weighted corrections based on operation history information such as recent interaction frequency and dwell time. The server encodes the sentiment results into a multi-dimensional sentiment state vector (e.g., happiness, sadness, anxiety, loneliness) and stores it in the sentiment state time series.

[0171] Output: An emotion state vector representing the user's current emotion.

[0172] Step 15: The server constructs a response and generates a prompt statement. The server generates prompts for dialogue responses based on the virtual object's personality attributes and the emotional state vector from step 14.

[0173] Input: Virtual object personality attributes, emotional state vector, current user messages and dialogue history summary.

[0174] The server fills this information into a dialogue template, generating a message such as: "You are a virtual avatar, quiet and gentle in personality, with a cat-like temperament, and good at listening to users. Your task is to accompany and comfort users, speaking softly and not too loudly. The user is currently feeling a bit sad, please respond gently." It then appends the user message and a brief dialogue history; simultaneously, it attaches numerical values ​​from the emotional state vector as control labels or embedding vectors for the generative AI model to use for style adjustment.

[0175] Output: Response prompts and related control vectors containing emotion control information.

[0176] Step 16: The server invokes a generative artificial intelligence model to generate dialogue responses. The server uses the text sub-model of a generative artificial intelligence model to reason about the prompt statement constructed in step 15 and generate the response content of the virtual object.

[0177] Input: Prompt statements and control vectors for response generation.

[0178] The server embeds the prompts as a vector sequence and concatenates emotion and personality control vectors into the model input. During decoding, the model adjusts the probability distribution based on the control vectors to make the output tone gentler and the rhythm more moderate. The server reconstructs the text from the word sequence output by the model, for example: "It sounds like you're not very happy today. I'm here with you. Do you want to tell me what happened? Don't rush, just tell me slowly." Output: The dialogue response text of the virtual object, and optional emoji or action tags (such as "soft_smile").

[0179] Step 17: The server generates reaction data and sends it to the terminal. The server generates a response data object based on the text response and action tag from step 16 and returns it to the terminal.

[0180] Input: reply text, emoji / action tag, current virtual object identifier.

[0181] The server packages this information into reaction data, serializes it into a response message, and sends it to the terminal through the network interface; at the same time, it appends a "virtual object speaking" entry to the session record and updates the dialogue history summary and emotional state time series.

[0182] Output: A response data message containing reply text and action instructions.

[0183] Step 18: The terminal displays the response and drives the virtual object's actions. The terminal receives the response data returned by the server and updates the appearance of the virtual objects in the interface.

[0184] Input: Response data message (response text, action tag, etc.).

[0185] The terminal displays the reply text as a new chat bubble, and calls the local text-to-speech engine to read the content aloud when necessary; the terminal parses action tags and triggers facial expression changes or action animations on images or 3D models (such as switching to a smiling emoji or playing a slight nodding action), so that the virtual object responds to the user in a way that matches the personality and emotional settings.

[0186] Output: The dialogue content and the dynamic expressions and actions of virtual objects displayed on the terminal.

[0187] Step 19: Customize the system by adding user-input prompts. As users continue using the device, they can input additional prompts to adjust the virtual object, such as: "Make this avatar quieter." or "I hope he tells more jokes." Input: Additional natural language text used to adjust the properties of virtual objects.

[0188] The terminal performs the same preprocessing as in step 1 on the text, marks it as a "re-customization request", and sends it to the server.

[0189] Output: Customization request data containing appended prompts and virtual object identifiers.

[0190] Step 20: The server updates the virtual object characteristic data based on the appended prompt. After receiving the customization request, the server reconstructs and regenerates the prompt statement based on the new prompt statement, and calls the generative artificial intelligence model to update the personality and behavioral attributes of the virtual object.

[0191] Input: Append prompt statement, existing virtual object feature data.

[0192] The server extracts new preference tags such as "quieter" and "tells more jokes," updates the personality attribute vector (e.g., reduces extroversion, increases humor), and inputs the updated vector and the original settings into the generative artificial intelligence model. The server guides the model to output new personality descriptions and behavioral parameters through prompts. The server replaces or expands the original feature data with the new descriptions and parameters and assigns a new version number.

[0193] Output: Updated version of virtual object feature data, and send the main changes (such as updated personality descriptions) to the terminal so that the terminal adopts the new settings in subsequent dialogues and displays.

[0194] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0195] The following technical shortcomings exist in existing online services and virtual character interaction technologies.

[0196] First, servers typically generate the appearance and dialogue of virtual characters based solely on simple rules or static configurations, lacking in-depth modeling of user behavior and preference data. This results in virtual characters exhibiting essentially the same appearance and interaction patterns across different users, failing to achieve high-granularity personalization at the system level. Second, the server-side utilization of natural language prompts is rather crude, generally simply forwarding user input directly to the generative model. It lacks the ability to automatically construct prompts based on user profiles, scene context, and product data, causing generative AI models to fail to fully utilize multi-source structured data, leading to a disconnect between the generated results and specific application scenarios.

[0197] Furthermore, in interactive scenarios such as virtual stores, existing systems often rely on simple front-end scripts for local recommendations and displays. The servers lack a comprehensive processing flow centered on generative AI models, integrating user profiling, product recommendation, dialogue generation, and virtual character behavior control. This makes overall optimization of the computing architecture and data flow paths difficult. Additionally, while some systems record user interaction logs, servers primarily use them for simple statistical analysis. There's a lack of a closed-loop optimization mechanism for continuously updating user profiles, prompt templates, and virtual character behavior strategies based on interaction history, hindering long-term interaction quality and resource utilization efficiency at the system level.

[0198] Therefore, how to introduce a unified technical solution on the server side that integrates user profile modeling, automatic generation of prompts, generative artificial intelligence model invocation, and virtual character control and product recommendation, so that the server can realize dynamic and evolvable personalized virtual character interaction processes based on multi-source data, thereby improving the overall computer technology in terms of computational efficiency, model utilization efficiency, and interaction effects, has become the technical problem that this invention needs to solve.

[0199] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.

[0200] In this invention, the server includes a functional module for acquiring user behavior and preference information and generating user profile information through machine learning; a functional module for automatically generating prompt statements that define the appearance and personality information of virtual characters based on the user profile information and calling a generative artificial intelligence model to generate corresponding feature information; a functional module for generating image information or stereoscopic display information representing virtual characters based on the feature information and sending it to the terminal for presentation in virtual space; a functional module for acquiring user operation information and dialogue information in virtual store space and generating recommendation information presented by virtual characters based on the user profile information and product information; a functional module for constructing dialogue prompt statements based on user dialogue information and user profile information and calling a generative artificial intelligence model to generate response content to control the voice and behavior of virtual characters; and a functional module for storing interaction history information between users and virtual characters and updating user profile information and recommendation information based on the interaction history information to continuously optimize the appearance information, personality information, and product recommendation behavior of virtual characters. This allows for the formation of an end-to-end data processing and interactive control process on the server side, centered on a generative artificial intelligence model and using prompt statements as a unified interface. This enables the computer system to achieve collaborative optimization among multi-source data fusion, model invocation strategies, virtual character presentation, and recommendation logic, thereby improving the quality of personalized generation and the utilization rate of computing resources, and fundamentally improving interactive information processing technology based on virtual characters.

[0201] "Information processing device" refers to an electronic device that includes a processor, memory, and communication interface, used to execute program code and perform calculation and control processing on input data, and can be implemented in the form of a server or other computing device.

[0202] "User behavior information" refers to time-related data generated by user operations in information systems or virtual spaces, including browsing history, click history, dwell time, purchase history, conversation content, and interface operation sequences.

[0203] "User preference information" refers to configuration and feedback data used to characterize users' long-term or short-term interests, including user-defined interest options, questionnaire responses, rating information, and interest tags inferred by the system.

[0204] "Attribute information" refers to abstract data used to represent a user's preference or behavioral characteristics in a certain interest category or feature dimension, including forms such as labels, weight values, probability distributions, or vector representations.

[0205] "User profile information" refers to a structured data set composed of multiple attribute information, used to comprehensively describe a user's interests, habits, consumption characteristics, and interaction characteristics. It can be represented as key-value pairs, vectors, or other data structures.

[0206] "Prompt statements" refer to instruction text written in natural language or structured text, used to provide task descriptions, contextual information, constraints, and output requirements to generative artificial intelligence models, guiding the models to generate expected output content.

[0207] "Generative artificial intelligence models" refer to data processing models trained using machine learning techniques, especially deep learning techniques, which are used to automatically generate text, images, or other content based on input data, including but not limited to language generation models and multimodal generation models.

[0208] "Virtual character" refers to a digital person or anthropomorphic image presented in virtual space for interaction with users. Its appearance and personality information can be configured and adjusted based on user data.

[0209] "Appearance information" refers to a set of parameters used to describe the visual presentation of a virtual character, including body shape, facial features, clothing, color style, props, and related information about two-dimensional images or three-dimensional models.

[0210] "Personality information" refers to a set of parameters used to describe the behavioral style and emotional tendencies of virtual characters during interaction, including friendliness, liveliness, formality, humor, professionalism, and corresponding behavioral rules.

[0211] "Feature information" refers to intermediate or configuration data generated by generative artificial intelligence models or other algorithms to comprehensively represent the appearance, personality, and related control parameters of a virtual character.

[0212] "Image information" refers to the data used to present two-dimensional visual content of virtual characters on a display device, including bitmap images, vector images, and related texture data and rendering parameters.

[0213] "Information for stereoscopic display" refers to 3D model data, skeletal data, animation data, and depth information used to present virtual characters in a 3D display environment or virtual reality environment.

[0214] "Terminal device" refers to a user-side device that communicates with an information processing device and is used to present virtual space and virtual characters to the user, including mobile terminals, desktop terminals, head-mounted display devices or other visual interactive devices.

[0215] "Virtual space" refers to a digital environment generated and presented on a terminal device through computer graphics and interactive technologies, including two-dimensional interface environments and three-dimensional immersive environments.

[0216] "Virtual store space" refers to a digital store environment built in a virtual space for displaying product information and providing product browsing and purchasing operations. It may include shelf areas, product display areas, and interactive areas.

[0217] "Product information" refers to structured data related to products to be recommended or displayed, including product identifiers, categories, attribute parameters, prices, inventory status, review information, and multimedia content links.

[0218] "Recommendation information" refers to the result data generated by the system based on user profile information and product information, used to indicate or rank one or more candidate products, including the list of recommended products, ranking weight, display priority, and reasons for recommendation.

[0219] "Dialogue information" refers to language-related data generated during the interaction between a user and a virtual character, including natural language text input by the user, text obtained by speech recognition, and responses and dialogue context output by the virtual character.

[0220] "Conversational information" refers to structured data used to drive virtual characters to perform language output and behavioral control, including generated response text, tone labels, action instructions, and contextual tags related to the current conversation state.

[0221] "Response content" refers to the semantic content generated or output by the virtual character in response to user input during the dialogue process, including answers, explanations, questions, suggestions, reminders, and emotional feedback.

[0222] "Voice and behavior" refers to the multimodal performance of virtual characters in virtual space related to the output content, including voice playback, lip movements, facial expressions, body movements, and changes in position and posture in the virtual scene.

[0223] "Interaction history information" refers to the time-series data records accumulated during the continuous interaction between users and virtual characters, including the content of each session, the sequence of user operations, feedback on recommendation results, and the resulting conversion behaviors.

[0224] The embodiments of the present invention will be described in detail below with reference to the appended notes. It should be understood that the present invention can be implemented in various specific ways, and the following description is merely exemplary and does not constitute a limitation on the technical scope of the present invention.

[0225] In this embodiment of the invention, the server, terminal, and user respectively assume different roles in computational processing, interactive presentation, and input requirements. The server, through specific data structures, algorithmic processes, and generative artificial intelligence model (GAI) invocation strategies, achieves personalized generation and control of virtual characters (hereinafter also referred to as "avatars"), and through presentation and interaction in the virtual store space, achieves the technical effect of improving the internal processing efficiency and generation quality of the computer system.

[0226] I. Example of System Hardware and Software Composition In one implementation, a server includes a multi-core central processing unit (CPU, such as a server-class processor based on a general-purpose instruction set), a graphics processing unit (GPU, such as an accelerator card supporting massively parallel matrix operations), main memory, non-volatile storage devices, and a network interface. The server runs an application framework (such as a backend framework based on an interpreted language) on top of an operating system (such as a UNIX-based server operating system) and runs the following software components: The server runs a database management system (such as a relational database system) to store user behavior records, preference data, product data, and user profile data; the server runs a distributed caching system to temporarily store frequently accessed data structures; the server runs a development framework for machine learning (such as a deep learning framework based on tensor computing) to implement interest recognition models and recommendation models; the server runs a generative artificial intelligence model inference client to call language generation models and image generation models deployed on local or external computing resources through a network interface.

[0227] In one embodiment, the terminal is a smartphone, tablet computing device, desktop computing device, or head-mounted display device. The terminal includes a processor, a graphics rendering unit, a display device, an input device (touchscreen, buttons, microphone, camera, etc.), and a communication interface. The terminal runs a virtual store application or browser front-end program on a mobile or desktop operating system. This program embeds a 3D rendering engine or graphics library to present the virtual store space and virtual characters.

[0228] Users operate through the terminal, including entering login information, selecting product categories, moving the viewpoint or position in the virtual space, and asking questions in natural language to the virtual character.

[0229] II. User Profile Construction and Data Structure In one implementation, the server maintains a user profile data structure for each user. The server defines this data structure as a record containing multiple fields, for example: The server stores the "behavioral feature field" as a multi-dimensional vector, with dimensions corresponding to product category, operation type, time window, etc.; the server stores the "preference label field" as several category labels and their weight values; the server stores the "price sensitivity field", "interaction activity field", etc. as numerical features; and the server uses the timestamp field to mark the most recent update time.

[0230] The server generates and updates the aforementioned user profiles using a machine learning model. In one implementation, the server employs a neural network model based on embedding layers and multiple fully connected layers. The server categorizes input feature vectors into discrete and continuous features. Discrete features (such as product category IDs and brand IDs) are mapped into low-dimensional dense vectors through the embedding layer, while continuous features (such as pageview counts and total spending) are standardized and directly input into subsequent network layers. The server uses a multilayer perceptron (MLP) structure with non-linear activation functions in the hidden layers, and generates probability values ​​for each interest tag using either the sigmoid or softmax function in the output layer.

[0231] During the training phase, the server uses historical user behavior data as a supervision signal, treating purchase behaviors and long-stay behaviors as positive samples and unselected behaviors as negative samples. The server uses binary cross-entropy or multi-class cross-entropy as the loss function. The server updates the network weights through backpropagation and gradient descent or its variants (such as the Adam optimization algorithm). During training, the server can employ data augmentation techniques, such as randomly perturbing the time window and pruning and reordering behavior sequences, to improve the model's robustness to changes in behavior patterns.

[0232] During the inference phase, the server receives current and historical user behavior data, encodes it into the aforementioned feature vectors, and inputs them into the interest recognition model. The server then obtains the probability distribution of each interest tag. For example, the server can obtain the probability values ​​corresponding to tags such as "pet-related," "outdoor activities," and "home appliances." Based on set thresholds and ranking results, the server generates a list of interest tags and their weights for the user profile, thus forming the user profile information.

[0233] Through the specific data and model structures described above, the server is able to compress and extract user behavior patterns in a lower-dimensional representation space. Compared to simple rule-based statistical methods, the server achieves improvements in inference speed and interest recognition accuracy, thereby providing more reliable basic features when generating prompts and controlling generative artificial intelligence models.

[0234] III. Generation of Avatar Setting Prompt Statements In one implementation, the server automatically constructs prompts for avatar setting based on user profile information. The server predefines several templates, which use natural language to describe task rules and generally include appearance descriptions, personality descriptions, and conversational style descriptions. The server then fills in placeholders in the templates with weighted interest tags from the user profile.

[0235] For example, when the tags "pet dog" and "outdoor equipment" have high weights in the user profile, the server generates the following prompt: "Users like dogs and are very interested in outdoor products. Please generate an avatar setting for the virtual shopping system:" 1. Appearance: Modeled after a dog, with a realistic style, wearing outdoor clothing such as hiking jackets and backpacks.

[0236] 2. Personality: Lively, humorous, patient, professional, and skilled at recommending products suitable for camping and hiking.

[0237] 3. Conversation style: Friendly yet professional tone, proactively inquiring about user needs and explaining technical parameters in easy-to-understand language.

[0238] Please use a structured description to output the avatar's appearance elements, personality tags, and dialogue style description. When constructing prompts, the server applies a series of rules that deviate from human habits. For example, the server sorts multiple tags by interest weight and selects only the top-ranked tags to control the length of the prompts; the server applies different sub-templates to different combinations of interests to control the diversity of results; and the server adds short-term interest fields based on the user's recent behavior, making the prompts more closely aligned with the user's current needs. These rules differ from simply concatenating user input into fixed text; they represent a template selection and filling strategy driven by machine learning results, which helps improve the efficiency of generative artificial intelligence models in utilizing contextual information.

[0239] The server sends the generated prompts to the inference server endpoint of the generative AI model via a network interface. In one implementation, the generative AI model can employ an autoregressive language model based on a multi-layered Transformer structure. The server specifies the model identifier, temperature parameters, maximum generation length, etc., when calling the model. The server receives the text output returned by the model and then parses the output into an internal avatar-based data structure, such as using key-value pairs to map appearance type, color style, clothing elements, personality traits, and dialogue style tags.

[0240] IV. Generation of Headshot Images and 3D Representation In one implementation, the server generates a visual representation based on a data structure set by the avatar. The server can employ two methods.

[0241] In one approach, the server invokes an image generation model, which can be a diffusion-based generation model. The server generates a new image-based tooltip based on the avatar settings, for example: "Please generate an image: a standing dog wearing an outdoor hiking jacket and backpack, in a realistic style, with an outdoor campsite as the background." The server inputs the prompt into the image generation model, and performs an iterative denoising process using a diffusion model on the GPU, ultimately obtaining a two-dimensional image of the head. The server stores this image in an object storage system and generates an access path for it.

[0242] In another approach, the server generates avatars based on a pre-built library of 3D characters. The server selects a base 3D skeleton and mesh model according to the category and style tags in the avatar settings. It maps appearance details (such as ear shape, fur color, and clothing) to parameters and adjusts the mesh and textures programmatically. The server generates a set of animation data, including basic actions such as walking, waving, turning, and pointing. Finally, the server combines the 3D model files, texture images, and animation configurations into an "avatar resource pack" and writes the location or download address of the resource pack into the system database.

[0243] Through the aforementioned image or 3D representation generation process, the server allows virtual character appearances to be dynamically generated through algorithmic combinations and generative models, rather than being limited to a fixed resource library. Because the server employs parametric modeling and generative model inference, it can significantly expand the available appearance combination space while maintaining controllable generation speed, thereby technically improving the system's generation diversity and storage efficiency.

[0244] V. Presentation and Interaction in Virtual Store Space In one implementation, the terminal requests avatar resource package information and avatar settings associated with the user from the server. The terminal receives the resource package address and setting data through a network interface. The terminal uses an embedded 3D rendering engine to load model files, textures, and animation configurations, and constructs a virtual store scene on its local graphics processing unit, including product shelves, display areas, and navigation areas.

[0245] The terminal places the virtual character in a designated location within the scene (e.g., near the entrance or the current product area) based on the avatar settings provided by the server. During initialization, the terminal determines whether to play a welcome animation and opening remarks based on personality information. During the rendering process, the terminal adjusts the level of detail of the avatar model according to frame rate scheduling and LOD (Level of Detail) strategies to reduce the graphics computation load while ensuring visual quality.

[0246] Users interact with virtual characters through touch controls on the terminal, controller controls, or head-mounted display gaze control, such as clicking on the virtual character, moving closer to the virtual character, or asking questions via voice input. After detecting the trigger condition, the terminal sends the user's input text question or the text data after speech recognition, along with contextual information (current browsing category, scene location, historical dialogue summary, etc.), to the server.

[0247] VI. Dialogue prompts and product recommendation generation After receiving the user's question and context, the server first retrieves a set of candidate products relevant to the current scenario from the user profile and product database. The server can then use another deep learning recommendation model (e.g., a click-through rate prediction model incorporating an attention mechanism) to input the user profile vector, the current context vector (such as the currently viewed category and price range), and the product feature vectors into the model. The server obtains a score for each candidate product. Based on these scores, the server selects a limited number of products as recommendation candidates.

[0248] The server then constructs dialogue prompts to guide the generative AI model in generating natural language responses. The server concatenates system role settings, user profile summaries, product candidate descriptions, and user questions into structured natural language text. For example: "You are a virtual shopping assistant avatar, which looks like a dog that loves outdoor sports, and has a lively, humorous, and professional personality."

[0249] User profile: Likes dogs, enjoys camping and hiking, and frequently buys outdoor equipment.

[0250] Current scenario: The user is browsing the camping equipment section of a virtual store.

[0251] User question: 'Are there any tents suitable for camping with a dog?' Product data (brief): - Tent A: Three-person tent with a vestibule, pet mat storage, waterproof rating 3000mm.

[0252] - Tent B: Two-person tent, lightweight, suitable for long-distance hiking, with a smaller front yard area.

[0253] - Tent C: A family tent for four people, with a large interior space and partitioned areas, including separate space for pets.

[0254] Based on user profiles and product data, please recommend 1-2 of the most suitable tents in a friendly, professional, and concise manner, and explain your reasons for the recommendations. When constructing the aforementioned prompt statements, the server executes the rules defined in the strategy module. These rules include controlling the length of the product data summary, selecting the attribute fields most relevant to the user's interests, and avoiding repeatedly displaying products that the user has explicitly rejected. These rules are applied uniformly in a programmatic manner based on the data structure and scoring results. Unlike manually editing fixed scripts, this is an algorithm-optimized prompt statement construction process.

[0255] The server inputs the prompt into the generative AI model, which internally models the context through a multi-layered self-attention mechanism and progressively generates the response text. The server receives this text and performs rule-based post-processing as necessary, such as filtering prohibited words, limiting the number of characters, and inserting clickable product identifiers. The result is then packaged into conversational information and sent to the terminal.

[0256] By combining a recommendation model with a generative AI model, the server technically achieves synergy between structured rating output and natural language generation: the recommendation model focuses on computational relevance and ranking accuracy, while the generative AI model handles language expression and interactive coherence. The server embeds the recommendation results into the natural language context through prompts, allowing the outputs of both models to be fused at a unified text interface layer.

[0257] VII. Virtual Character Behavior Control and Multimodal Output After receiving the dialogue information from the server, the terminal displays its response text as a speech bubble on the interface. Simultaneously, the terminal uses a local or cloud-based text-to-speech module to generate an audio signal. The terminal then controls the playback timing of the audio based on speech rate, pauses, and intonation tags.

[0258] The terminal drives the virtual character to execute corresponding animations based on behavioral instructions in the response content (such as "point to tent A", "approach the user", "nod"). In its rendering logic, the terminal aligns the action sequence with the voice playback, thereby achieving lip-syncing and facial expression interaction. On the product display interface, the terminal visually emphasizes recommended products through highlighting, zooming, or repositioning, achieving synchronization between the virtual character's behavior and interface changes.

[0259] After hearing and seeing the introduction of the virtual character, users can click on recommended products to view details or add them to their shopping cart. The terminal sends these action events back to the server, which records them as part of the interaction history.

[0260] By realizing the combined presentation of voice, image and 3D motion on the terminal side, this invention not only outputs text results, but also transforms the dialogue information generated on the server side into a multimodal representation that fits the scene, thereby providing higher information density and understanding efficiency at the user interaction level.

[0261] 8. Continuous Optimization Driven by Interaction History In one implementation, the server continuously collects and analyzes interaction history information. The server stores user questions, virtual character responses, recommended product lists, and subsequent user behaviors (clicks, favorites, purchases, feedback of disinterest, etc.) in each session as structured logs, including timestamps, session identifiers, user identifiers, and context tags.

[0262] The server periodically analyzes these logs using a batch computing framework. When training a new recommendation model, the server uses whether a user accepts a recommendation (e.g., clicks or purchases) as a supervisory signal, and the user profile vector and product feature vector used in the recommendation process as model inputs, training with cross-entropy or other suitable loss functions. At the suggestion strategy level, the server can also use these logs to statistically analyze the conversion rate differences resulting from different templates and information combinations, retaining and prioritizing efficient templates to progressively optimize the suggestion generation strategy.

[0263] Through the aforementioned closed-loop mechanism, the server continuously updates the interest weights in the user profile, the recommendation model parameters, and the prompt templates, enabling the entire system to automatically adapt to changes in the user group and product structure during long-term operation. From a computational perspective, this design not only improves algorithm accuracy but also distributes the computational load by combining online and offline methods, achieving lightweight inference and centralized training, thereby improving overall system processing efficiency and reducing real-time communication burden.

[0264] IX. Technical Effects and Improvements This invention introduces an integrated architecture on the server side, encompassing user profile modeling, automatic generation of prompts, invocation of generative artificial intelligence models, recommendation model inference, and virtual character control. This architecture enables computer systems to handle complex human-computer interaction tasks with a unified data flow and modular division. Instead of simply forwarding user input to a single model, the server effectively integrates multi-source data into the generation process through specific feature extraction, model structure, data structure, and prompt strategy.

[0265] Because the server employs vectorized representations, embedding layers, and multi-layer neural networks, it can perform batch inference with high throughput in the interest identification and recommendation stages, reducing the average computation time required for each interaction. By automatically generating prompts and compressing structured information into more information-dense natural language context, the server reduces the number of round trips and redundant data transmissions with generative AI models, thus alleviating communication load. Furthermore, by referencing user profiles and interaction history when generating avatars and dialogue content, the server reduces the need for users to repeatedly input their preferences, lowering the error rate and the proportion of invalid dialogues during human-computer interaction.

[0266] Furthermore, this invention reduces the reliance on storing a large number of fixed resource files through parametric avatar generation and 3D resource combination. The server can generate or combine avatars on demand, significantly reducing resource management complexity and storage space consumption. These improvements are reflected in the data structure design, algorithm flow, and resource management strategies within the computer system, and are not merely automation at the business rule level.

[0267] 10. Other forms of implementation In an alternative implementation, the server can replace the interest recognition and recommendation models with sequence model structures, such as models based on bidirectional encoders and self-attention mechanisms, to more precisely capture the sequential dependencies in the behavioral sequences. In this case, the server adjusts the feature input format, changing from static aggregated features to time-step sequence features, to further improve the response speed to short-term interest changes.

[0268] In another implementation, the server can deploy a streamlined version of the generative AI model locally to provide rapid response in low-latency scenarios, while calling a remote, high-capacity model when complex generation or high-quality text is required. In this multi-model collaborative architecture, the server uses a routing strategy module to select the appropriate model based on problem complexity, user level, or network conditions, further optimizing the balance between response time and generation quality at the system level.

[0269] In another implementation, the terminal can have the ability to locally cache avatar resources. After initially loading the avatar resource package, the terminal stores it locally. In subsequent sessions, it only needs to request incremental updates from the server (such as changes in personality parameters or conversation style), reducing the bandwidth overhead of repeated downloads. This approach, combining server-side parameter updates with terminal-side resource reuse, can effectively reduce network load and improve page and scene loading speeds.

[0270] Through the above-mentioned various implementation forms, the system of the present invention can be flexibly deployed in different hardware environments and business scenarios. At the same time, it always revolves around the collaborative processing mechanism of user profiles, prompt statements and generative artificial intelligence models, thereby achieving substantial improvements in the internal information processing flow and resource utilization of computers.

[0271] use Figure 12 The processing flow is explained.

[0272] Step 1: The server obtains basic user information and historical behavior data.

[0273] Input: User ID and session token uploaded by the terminal, as well as raw data tables stored in the server database, such as browsing history, purchase history, questionnaire results, and collection history.

[0274] The server queries multiple data tables related to the user in the database based on the user identifier, and reads records containing fields such as product ID, timestamp, operation type, page dwell time, and questionnaire options.

[0275] The server deduplicates, sorts by time, and fills in missing values ​​on the read records, assembling the results into a time-ordered sequence of behaviors and a set of preference configurations.

[0276] Output: Structured raw user behavior sequence data and preference configuration data.

[0277] Step 2: The server generates user feature vectors and constructs user profile information.

[0278] Input: Behavioral sequence data and preference configuration data obtained in step 1.

[0279] The server uses a feature engineering module to encode discrete features (product category, brand, operation type, etc.) and map them to integer IDs or embedded indexes; the server normalizes continuous features (number of times, amount, duration, etc.).

[0280] The server concatenates the processed features into fixed-length or variable-length feature vectors, which serve as the input tensors for the machine learning model.

[0281] The server calls a pre-trained interest recognition neural network in the deep learning framework to perform forward inference on the input tensor and obtain the probability or weight values ​​corresponding to multiple interest labels.

[0282] The server filters out key interest tags based on set thresholds and sorting rules, and encapsulates the tags and their weights, along with other statistical features, into user profile information.

[0283] Output: A user profile data structure containing interest tags, weights, and preference statistical features.

[0284] Step 3: The server generates profile pictures and sets prompts based on user profiles.

[0285] Input: User profile information generated in step 2 (interest tags and weights, recent behavior summary, etc.).

[0286] The server sorts the tags according to their interest tag weights and selects several tags with higher weights, such as "pet dogs" and "outdoor products".

[0287] The server selects a prompt template that corresponds to the tag combination and fills the template placeholders with information such as interest tags, personality traits, and conversation style.

[0288] The server generates a prompt message in natural language for setting the avatar, for example: "Users like dogs and are very interested in outdoor products. Please generate an avatar setting for the virtual shopping system:" 1. Appearance: Modeled after a dog, with a realistic style, wearing outdoor clothing such as hiking jackets and backpacks.

[0289] 2. Personality: Lively, humorous, patient, professional, and skilled at recommending products suitable for camping and hiking.

[0290] 3. Conversation style: Friendly yet professional tone, proactively inquiring about user needs and explaining technical parameters in easy-to-understand language.

[0291] Please use a structured description to output the avatar's appearance elements, personality tags, and dialogue style description. The server performs length checks and sensitive word filtering on the prompt statements.

[0292] Output: A text-based prompt for setting your profile picture.

[0293] Step 4: The server calls a generative artificial intelligence model to generate avatars and set feature information.

[0294] Input: The avatar setting prompt obtained in step 3.

[0295] The server sends a request to the generative artificial intelligence model inference service via a network interface, including a prompt statement as part of the request body, along with parameters such as model identifier, maximum generation length, and temperature.

[0296] The server waits for the model to return the generated result text, and after receiving the response, it parses the text and extracts the appearance elements (image type, clothing style, color scheme, etc.), personality tags (lively, humorous, professional, etc.), and dialogue style (formality, word choice, etc.) described in it into structured fields.

[0297] The server organizes these fields into an avatar feature information data structure, which includes appearance information, personality information, and conversational style information.

[0298] Output: Avatar feature information used to control the appearance and personality of the avatar.

[0299] Step 5: The server generates images or 3D representations of the avatar.

[0300] Input: The avatar feature information obtained in step 4.

[0301] The server constructs images based on the avatar's appearance information and generates prompts that describe visual elements such as the character's appearance, clothing, and scene style.

[0302] The server calls the image generation model or 3D resource combination module, sends prompts and resolution parameters to the image generation model, or selects a base model from the 3D resource library and applies appearance parameters.

[0303] The server performs image generation or 3D model parameter adjustment calculations on the GPU, generating 2D image files or 3D model files of the avatar, along with their textures and animation configurations.

[0304] The server stores the generated visual resources in a file storage system and records the access path and resource identifier.

[0305] Output: Portrait resource information including resource path, model file name, texture and animation configuration.

[0306] Step 6: The server sends avatar settings and resource information to the terminal, and the terminal loads and displays the avatar.

[0307] Input: Avatar feature information from step 4 and avatar resource information from step 5, as well as a request from the terminal (including user identifier and session information).

[0308] The server queries the corresponding avatar feature information and resource path based on the user identifier and assembles them into a response data packet.

[0309] The server sends the response data packet to the terminal through the network interface.

[0310] After receiving the data packet, the terminal loads the image file or 3D model file from the server or object storage and parses the model and texture in the local graphics engine.

[0311] The terminal initializes the avatar's posture and expression based on the avatar's personality information, and renders the avatar to the designated position in the virtual store scene.

[0312] Output: The visual avatar displayed on the terminal and the avatar instance object in memory.

[0313] Step 7: Users interact with avatars using natural language through the terminal.

[0314] Input: The avatar interface and interactive controls displayed on the terminal.

[0315] Users can click on their avatar, click the consultation button, or ask questions via voice input on the terminal, such as "Are there any tents suitable for camping with dogs?"

[0316] The terminal packages the user's text input or the text converted through speech recognition together with the current scene information (category, view position, list of viewed products) into a dialogue request.

[0317] The terminal sends the dialogue request to the server via network protocol.

[0318] Output: A dialogue request message containing the user's question text and context data.

[0319] Step 8: The server generates dialog prompts based on user profiles and context information.

[0320] Input: The dialogue request message (user question text, scenario information) in step 7, as well as the user profile information and relevant product records in the product database in step 2.

[0321] The server retrieves candidate products from the product database based on the current scenario and filters out a set of products that are relevant to the current category and user interests.

[0322] The server applies a recommendation model to score the candidate product set, using user profile vectors, product feature vectors, and contextual features as inputs, and outputs a relevance score for each product.

[0323] The server selects the top-scoring products, extracts their key attributes (name, specifications, applicable scenarios, etc.), and generates brief product description text.

[0324] The server combines system role settings, user profile overviews, scenario descriptions, user questions, and product data to construct dialog prompts for generative artificial intelligence models, such as: "You are a virtual shopping assistant avatar, which looks like a dog that loves outdoor sports, and has a lively, humorous, and professional personality."

[0325] User profile: Likes dogs, enjoys camping and hiking, and frequently buys outdoor equipment.

[0326] Current scenario: The user is browsing the camping equipment section of a virtual store.

[0327] User question: 'Are there any tents suitable for camping with a dog?' Product data (brief): - Tent A: Three-person tent with a vestibule, pet mat storage, waterproof rating 3000mm.

[0328] - Tent B: Two-person tent, lightweight, suitable for long-distance hiking, with a smaller front yard area.

[0329] - Tent C: A family tent for four people, with a large interior space and partitioned areas, including separate space for pets.

[0330] Based on user profiles and product data, please recommend 1-2 of the most suitable tents in a friendly, professional, and concise manner, and explain your reasons for the recommendations. The server performs length control and content filtering on the prompt statements.

[0331] Output: Dialogue prompts containing user questions, product summaries, and role settings.

[0332] Step 9: The server invokes a generative artificial intelligence model to generate response content.

[0333] Input: Prompt statements for the dialogue generated in step 8.

[0334] The server sends a request to the generative artificial intelligence model service via a network interface, using dialogue prompts as input text.

[0335] The server sets temperature, sampling strategy, and maximum output length in the generation parameters to control the diversity and length of the generated results.

[0336] The server receives natural language responses from generative artificial intelligence models, including recommended product names, reasons for recommendation, and supplementary information.

[0337] The server performs post-processing on the response text, such as detecting whether it contains sensitive content, trimming excessively long paragraphs, and inserting internal product identifiers.

[0338] Output: The processed response text.

[0339] Step 10: The server assembles the conversational information and recommendation data and sends them to the terminal.

[0340] Input: The response text from step 9, and the recommended product ID and related attribute data obtained in step 8.

[0341] The server merges the response text with the list of recommended products to construct a dialog message data structure, which includes the response text, the identifier of the recommended products, and the corresponding display parameters.

[0342] The server sends conversational information to the terminal via a network interface.

[0343] After receiving the dialogue information, the terminal displays the response text in the dialogue area of ​​the interface and calls the text-to-speech module to generate speech; at the same time, the terminal highlights the location of recommended products or marks them with labels such as "recommended" in the virtual store scene.

[0344] The terminal drives the avatar to perform actions such as pointing, nodding, and moving based on the behavioral instructions in the dialogue information.

[0345] Output: The dialogue responses and highlighted recommended products displayed on the terminal, as well as the avatars of those in interactive states.

[0346] Step 11: Users provide feedback on the recommended results, and the terminal and server record the interaction history.

[0347] Input: The list of recommended products displayed on the terminal interface and the dialogue content.

[0348] Users can click on a recommended product on the terminal to view details, add it to their shopping cart, or complete the purchase. They can also ignore the recommendation or close the dialog box.

[0349] The terminal records the user's action type (click, purchase, ignore, etc.), action time, and session identifier for each recommended product as event data and sends it to the server.

[0350] The server receives event data and writes it to the interaction history log, which includes fields such as user ID, session ID, recommendation list, user behavior, and timestamp.

[0351] The server may optionally incrementally update user profiles based on the latest interaction history, adjust interest weights, or record the degree of acceptance of specific product categories.

[0352] Output: Updated interaction history information and potentially updated user profile information.

[0353] Step 12: The server uses an interaction history optimization model and a prompt statement generation strategy.

[0354] Input: The accumulated interaction history log from step 11.

[0355] In an offline environment, the server uses a batch processing computing framework to load a large amount of historical interaction data. Whether the recommendation result was clicked or purchased is used as a supervision label, and the user profile vector and product feature vector used in the recommendation are used as input features to retrain or fine-tune the parameters of the recommendation model.

[0356] The server analyzes the conversion effects of different types of prompt templates and information combinations, identifies templates with higher click-through rates or conversion rates, and prioritizes the use of these templates in the online environment.

[0357] The server deploys the updated model parameters and template configuration to the online inference service, making subsequent recommendation scoring and prompt statement construction more consistent with user behavior patterns.

[0358] Output: Updated recommendation model, prompt statement template priority configuration, and the latest set of parameters for online services.

[0359] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.

[0360] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."

[0361] In existing technologies, computer systems for personal health management and daily life management typically provide users with uniform prompts through simple rules or fixed templates, making it difficult to reflect users' physiological states and daily routine changes in a timely and accurate manner. Furthermore, existing systems often only perform data storage and basic statistics on the server side, lacking a collaborative processing mechanism that can uniformly clean and analyze multi-source time-series health data on the server side, construct high-quality prompts based on the analysis results, and drive generative artificial intelligence models to output personalized text suggestions.

[0362] Furthermore, most existing systems that utilize virtual avatars for interaction only switch avatar expressions and dialogue on the terminal side based on preset scripts or limited emotion tags. The correlation between their emotional responses and the user's actual health status and behavioral feedback is weak, making it difficult to form a closed-loop optimization. In addition, generative artificial intelligence models are often simply regarded as "text generation black boxes." The system lacks the technical means to dynamically adjust the design of prompts and model parameters based on user feedback, resulting in the quality of suggestions remaining at a fixed level for a long time, failing to reflect the performance improvement of computer systems that "continuously improve with use."

[0363] Furthermore, in a typical cloud-edge architecture, the data flow between the server and the terminal often revolves only around uploading and downloading content, lacking a dedicated cloud-edge collaboration process designed for "health management + schedule management + virtual avatar interaction": including how to organize and manage health-related information, schedule information, user feedback information, and generative artificial intelligence model call results on the server side; how to unify these heterogeneous information into prompt statements that can be efficiently utilized by the model on the server side; and how to control the avatar behavior generation logic on the server side, so that the terminal only needs to perform lightweight display and broadcast processing, thereby improving the overall system's processing efficiency and scalability under large-scale user concurrency.

[0364] In summary, it is necessary to provide a new computer implementation approach to achieve integrated data processing and generative AI-driven interactive control for health management and daily life management on the server side. On the one hand, by centrally analyzing time-series physiological indicators and constructing structured analysis results and prompts, the relevance and consistency of generated suggestions can be fundamentally improved. On the other hand, by unifying the management of avatar display data generation, emotion inference, and feedback-driven model adaptive updates, the algorithm performance and interactive experience of the entire system can continuously evolve on the server side, thereby substantially improving the data processing efficiency and intelligence level of the computer system in this application scenario.

[0365] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.

[0366] In this invention, the server includes: a device for instructing a terminal to acquire health-related information, including user physiological information, and schedule information, and receiving health-related information and schedule information from the terminal; a device for processing data based on health-related information to generate analysis results about the user's health status and lifestyle habits; a device for constructing prompt statements as input to a generative artificial intelligence model based on the analysis results and schedule information, and inputting the prompt statements and analysis results into the generative artificial intelligence model so that the generative artificial intelligence model generates response text containing health management suggestions and daily life suggestions for different users; a device for generating avatar display data as avatar voice and action content from the response text, and sending the avatar display data to the terminal so that the terminal can visually display the avatar; a device for updating the prompt statements input to the generative artificial intelligence model and / or the operating parameters of the generative artificial intelligence model based on feedback information from the user obtained through the terminal, thereby gradually improving the content of health management suggestions and daily life suggestions; and a device for inferring the user's emotional state based on health-related information, schedule information, and feedback information, constructing prompt statements for causing the generative artificial intelligence model to generate avatar response content corresponding to the emotional state, and sending the avatar response content to the terminal. This allows for a centralized analysis and generation control process on the server side, focusing on health and schedule data. By standardizing prompt statements and model invocation mechanisms, heterogeneous data is unified into inputs that can be efficiently processed by generative artificial intelligence models, thereby improving the quality and consistency of suggestion generation. Simultaneously, through feedback-driven prompt statements and adaptive updates of model parameters, the system's reasoning capabilities and interaction strategies are continuously optimized on the server side, reducing the terminal processing burden, improving overall computing resource utilization and system scalability, and achieving substantial improvements to existing computer technology in the fields of health management and virtual avatar interaction.

[0367] A “system” refers to a collection of devices including at least one server, at least one terminal, and computer programs running on them for implementing health management and daily life management functions, which interact and process data collaboratively through a network.

[0368] A "server" is an information processing device that has a processor, memory, and network interface, used to centrally process health-related information, schedule information, and feedback information from terminals, and to perform functions such as data analysis, generative artificial intelligence model invocation, and avatar display data generation.

[0369] "Terminal" refers to a user-side information processing device operated by the user and communicating with the server to collect user data, display avatars and suggested content, and receive reminders and push information, including but not limited to mobile communication terminals, wearable terminals and fixed terminals.

[0370] "Health-related information" refers to a set of data related to a user's physiological state and health status, including time-series physiological indicators such as heart rate, steps, and body temperature, and their derived data collected by wearable information processing devices or other sensors.

[0371] "Schedule information" refers to time-schedule data that indicates events, tasks, or reminders that a user needs to perform at a specific time or time period, including medication times, exercise plans, work or life schedules, and their repetition rules.

[0372] "Physiological information" refers to measurement data that reflects the user's physical condition, including numerical information that can be obtained directly or indirectly by sensors, such as heart rate, respiratory rate, steps, body temperature, and activity level.

[0373] "Time series physiological indicators" refer to a series of physiological information data collected and recorded continuously in chronological order, used to reflect the physiological change trends of users at different points in time or over different time periods.

[0374] "Data processing" refers to the calculations and processing of received health-related and schedule information on the server, including data cleaning, format conversion, statistical calculation, outlier detection, aggregation, classification, and feature extraction.

[0375] "Analysis results" refers to the structured information obtained by the server based on data processing, which is used to characterize the user's health status and lifestyle characteristics, including statistical indicators, risk assessment, activity level classification, and trend judgment.

[0376] "Generative artificial intelligence models" refer to artificial intelligence models trained through machine learning or deep learning methods that can automatically generate natural language text or other forms of output based on input prompts and structured data. These include language models, dialogue models, and multimodal generative models.

[0377] "Prompt statements" refer to natural language or semi-structured text used as input to generative artificial intelligence models. They describe the user's basic attributes, analysis results, schedule information, and specify the word count range, expression style, and number of specific action items in the output content to guide the generative artificial intelligence model in generating expected response text.

[0378] "Response text" refers to the natural language text content output by the generative artificial intelligence model after receiving prompts and related data, which includes at least health management suggestions and / or daily life suggestions for the user.

[0379] "Health management suggestions" refer to natural language suggestions generated based on user health-related information and analysis results, used to guide users to adjust their exercise, rest, diet and other behaviors to improve or maintain their health status.

[0380] "Daily life suggestions" refer to natural language suggestions generated based on users' schedule information and behavior patterns, related to time management, task priorities, and activity arrangements, used to optimize users' daily life management.

[0381] "Avatar" refers to a virtual character or image presented on a terminal in the form of images, animations, 3D models, etc., used to interact with users and carry response text and emotional content.

[0382] "Avatar display data" refers to the control data used to drive the avatar display on the terminal, including the avatar's appearance information, facial expression information, action information, and time control information associated with the voice content.

[0383] "Speech content" refers to the speech text or speech signal converted from the response text or avatar response content and used for output in speech form.

[0384] "Action content" refers to the postures, facial expressions, body movements, or other animated behaviors performed by the avatar during the display process, used in conjunction with spoken and textual content.

[0385] "Feedback information" refers to information such as user evaluations, operation records, or subjective opinions on response text, avatar performance, or system functions on the terminal, including ratings, text comments, whether suggestions are adopted, and interaction frequency.

[0386] "Running parameters" refer to adjustable variables used to control the internal inference process and output characteristics of generative artificial intelligence models, including but not limited to temperature parameters, sampling strategy parameters, output length limits, and specific task weights.

[0387] "Emotional state" refers to the category or quantitative indicator of a user's psychological or emotional tendency, inferred from the user's health-related information, schedule information, interaction behavior and feedback information, such as tension, relaxation, fatigue, positivity, etc.

[0388] "Avatar response content" refers to natural language text, vocal content, and accompanying action content generated by a generative artificial intelligence model in response to an inferred emotional state, used to express responses such as comfort, encouragement, reminders, or empathy.

[0389] This invention will describe in detail the hardware configuration, software modules, data structures, generative artificial intelligence models, and prompt statement design of the system, incorporating collaborative processing between the server, terminal, and user. The following description is merely an exemplary embodiment, and the invention is not limited thereto.

[0390] First, let me explain the overall structure.

[0391] A server can consist of a general-purpose information processing device with a multi-core central processing unit, main memory, large-capacity non-volatile storage, and a network interface. The server can run a general-purpose server operating system, such as a kernel-based server operating system, and deploy a container runtime environment (e.g., a container platform) on top of it, running backend service programs within the containers. The server can install scripting language runtime environments (e.g., Python 3.x), statistical analysis environments (e.g., R4.x), and use relational database management systems (e.g., general-purpose relational databases) or document-oriented database systems to store user and health data.

[0392] The terminal can be a smart device running a mobile operating system, such as a mobile phone or tablet. The terminal can run native applications developed using programming languages ​​(such as Kotlin / Java or Swift), or applications developed using cross-platform frameworks. The terminal may include a Bluetooth communication module, a wireless LAN module, a local storage module (such as an SQLite database), a graphics processing unit, and an audio output module. The terminal can establish a connection with wearable information processing devices (smart bracelets, smartwatches, etc.) via the Bluetooth Low Energy communication protocol to collect physiological data such as heart rate, steps, and body temperature.

[0393] Users wear wearable information processing devices in real-world environments and interact with servers through applications installed on the terminals.

[0394] I. Server program generation and deployment module structure The server can deploy multiple functional modules in a container environment, including: a data receiving module, a data preprocessing module, a health analysis module, a schedule management module, a generative artificial intelligence interface module, a prompt statement construction module, an avatar display data generation module, a feedback processing module, and a model adaptive optimization module.

[0395] The server can use a web application framework (such as a Python-based web framework or a script-based web framework) to implement a terminal-oriented application interface, which communicates with the terminal via the secure transport protocol (HTTPS).

[0396] During deployment, the server can generate and store program configuration files, which include parameters such as the service addresses of each module, the endpoint addresses of the generative AI model, the model version number, prompt statement templates, data sampling intervals, and the length of the analysis time window. By centrally managing these parameters on the server side, system behavior can be adjusted without modifying the terminal program, thereby improving system scalability and maintainability.

[0397] II. Server performs data processing and calculations. 1. Data structure and storage of health-related information The server can store health-related information uploaded by the terminal in a health record table. The server can use a relational database, creating a table such as `health_records` with fields including: - user_id: User identifier; - timestamp: timestamp; - heart_rate: Heart rate value; - steps: Current cumulative steps; - Temperature: Body temperature or skin temperature; - device_id: Identifier for wearable device.

[0398] The server can create a composite index on the user_id and timestamp fields of the health_records table to enable fast retrieval by user and time range. This data structure design allows the server to efficiently retrieve physiological data over continuous time periods, providing a foundation for subsequent time-series analysis.

[0399] 2. Data Preprocessing and Feature Calculation The server can use Python data analysis libraries (such as Pandas and NumPy) or statistical analysis environments in the data preprocessing module to clean and calculate features of health-related information. The server can perform the following technical processing: The server can detect outliers in heart rate data. For example, if the heart rate is below a preset lower limit or above a preset upper limit, the record will be marked as abnormal and removed. The server can use sliding window averaging or median filtering to smooth out short-term fluctuations and reduce the impact of noise on the analysis results.

[0400] The server can aggregate raw data in fixed time windows (e.g., 1 minute, 5 minutes), calculating the average heart rate, maximum heart rate, minimum heart rate, step increment, and average temperature within each window. This aggregated data can be stored in the `health_summary` table, with fields including `user_id`, `time_window_start`, `avg_heart_rate`, `max_heart_rate`, `min_heart_rate`, `step_delta`, and `avg_temperature`. By storing aggregated data, the server can significantly reduce the amount of data needed for subsequent analysis, improving computational efficiency.

[0401] The server can derive features from time-series data during the preprocessing stage, such as: resting heart rate estimation (taking low heart rate values ​​for a period of time at night or in the early morning), continuous sedentary time (statistically counting time periods where the number of steps changes close to zero), and daytime activity ratio. These features can serve as input features for generative artificial intelligence models, enabling the models to make more refined judgments.

[0402] 3. Algorithm Analysis of Health Status and Activity Level The server can use machine learning algorithms in the health analysis module to further analyze the preprocessed features. For example, the server can use algorithms such as logistic regression, support vector machines, decision trees, or gradient boosting trees to classify users' activity levels into categories such as "low activity," "moderate activity," and "high activity." The server can also determine possible fatigue or potential stress states based on resting heart rate trends and body temperature changes.

[0403] The server can pre-train these traditional machine learning models using labeled data on the server side, generating model parameter files (such as weight coefficients and tree structures) through offline training, and then loading and using them in the online system. By centralizing model inference on the server side, it is possible to avoid deploying complex algorithms on the terminal side, which helps to reduce the computational burden on the terminal.

[0404] The server generates "analysis results" based on the algorithm described above, represented by structured fields such as activity_level, resting_hr_trend, sedentary_hours, and weekly_step_avg. These fields will be used to construct prompts and invoke generative artificial intelligence models.

[0405] III. The server constructs prompt statements and invokes the generative artificial intelligence model. 1. Structure and Deployment of Generative Artificial Intelligence Models The server can deploy generative AI models on a separate model service node. This model can employ an attention-based, multi-layered deep neural network architecture, such as a multi-layered Transformer architecture, including: - Word embedding layer: Converts input Chinese characters or words into vector representations; - Multi-head self-attention layer: Extracts long-range dependency information; - Feedforward network layer: performs nonlinear transformations; - Layer normalization and residual connection module: stabilizes the training and inference process; - Output generation layer: Generates the next token by sampling step by step based on the probability distribution.

[0406] The server can train the model using deep learning frameworks such as PyTorch and TensorFlow. During training, the server can use cross-entropy as the loss function to minimize the difference between the generated text and the manually written reference text. The server can use adaptive moment estimation optimization algorithms to update parameters and iteratively update weights through batch gradient descent. The server can also expand the training sample set by recombining health data descriptions and suggestion texts from different user scenarios through data augmentation, thereby improving the model's generalization ability in various scenarios.

[0407] 2. Construction and function of prompt statements In the prompt statement construction module, the server can combine user basic attributes, analysis results, and schedule information into a natural language description and attach explicit output constraint instructions to form a prompt statement.

[0408] The server can use a template-based approach to construct prompt statements, examples of which include: The server can construct the following exercise suggestion prompts: "You are a professional health management consultant. Below is a user's basic information and health analysis results for the past 7 days. Please generate an exercise suggestion for the user that is no more than 200 words based on this information. Requirements:" 1) Use plain and easy-to-understand Chinese; 2) Provide suggested daily exercise types, durations, and approximate intensity; 3) Use a gentle and encouraging tone, and avoid using threatening language.

[0409] User basic information: Male, 35 years old, height 170cm, weight 78kg.

[0410] Health analysis results: Average daily steps: 4500; prolonged sitting time; resting heart rate slightly higher than the average for peers. The server can also generate comprehensive health summary and prompt statements: Please write a comprehensive health summary report for the user this week based on the following data, keeping it under 300 words. The report should include: an overall assessment of the week, key issues, and actionable recommendations for next week.

[0411] Health data summary: This week's average resting heart rate increased by 5 beats per minute compared to last week; average daily steps: 4800; daily sedentary time: more than 8 hours.

[0412] Please use a friendly and encouraging tone, and avoid using medical diagnostic terminology. By using this design of prompts that include "explanatory text + output constraint instructions", the server can impose constraints on the length, style, and structure of the output content at the model input level, making the generative artificial intelligence model more in line with system requirements when outputting, thereby reducing server-side post-processing overhead and improving overall processing efficiency.

[0413] 3. Model Invocation and Result Post-processing The server can concatenate the constructed prompt statements with the structured analysis results into an input sequence, which is then sent to the model service via HTTP or a remote procedure call interface. The model service performs forward inference on the GPU, generating response text token by token based on the prompt statements.

[0414] After receiving the response text, the server can perform post-processing steps, including sensitive word filtering, length truncation, format standardization, and adding disclaimers, to ensure the security and readability of the output. The final response text, presented as health management or daily life advice, is then sent to the avatar display data generation module.

[0415] IV. The server generates avatar display data and sends it to the terminal. In the avatar display data generation module, the server can convert the response text into the control data required for avatar display. The server can assign corresponding voice and action tags to each sentence or phrase based on text content, sentence boundaries, and emotional tone, for example: - Segment the text into a sequence of sentences; - Assign speech rate and intonation tags to each sentence; - Specify an expression curve for the overall content (e.g., a default smile, raised eyebrows for emphasis); - Assign specific actions (such as nodding or spreading gestures) to specific keywords (such as "relax" or "rest").

[0416] The server can encode the above control information into structured avatar display data, such as including: - text: Response text; - speech_profile: Speech parameters; - motion_script: motion sequences and timeline; - emotion_tag: Overall emotion tag.

[0417] The server returns the avatar display data to the terminal via API. By generating avatar control scripts uniformly on the server side, the terminal only needs to execute simple rendering and playback logic, thereby reducing the terminal's computational load and ensuring consistent avatar behavior across multiple terminals.

[0418] V. Specific Functions and Examples of Terminals The terminal can run a health management application locally, which is responsible for communicating with wearable devices, uploading data to the server, and displaying avatar data.

[0419] The terminal can use the Bluetooth interface provided by the system to read information such as heart rate, steps, and body temperature from wearable devices, temporarily store the data in a local database, and then upload it to the server in batches according to time.

[0420] After receiving the avatar display data returned by the server, the terminal can call the local animation engine or graphics library to render the avatar model on the screen and drive the changes in expressions and movements according to motion_script; the terminal can also call the text-to-speech interface to convert the text field into speech and play it in sync with the avatar's lip movements.

[0421] For example, when the server returns the suggestion "Today we recommend that you take a 20-minute light walk and get up and move around for 5 minutes every hour", the terminal can display a smiling virtual character and read the content slowly by voice, accompanied by nodding and gestures.

[0422] VI. Server Adaptive Optimization Based on Feedback Information The server can collect user feedback information in the feedback processing module. After displaying suggestions, the terminal can provide a "helpful / not helpful" button and a rating component; user actions on the terminal will generate evaluation data.

[0423] The server can associate and store feedback information with corresponding prompts, analysis results, and response text. The server can periodically use data analysis tools to calculate average scores, acceptance rates, bounce rates, and other metrics for different prompt templates.

[0424] The server can automatically adjust the prompt templates based on statistical results. For example, it can add constraints such as "Please increase the number of specific action suggestions" or "Please keep it under 150 characters" to templates with low scores, thereby optimizing the behavior of the generative AI model at the prompt level. This approach, which iterates through prompt templates rather than relying solely on offline retraining, can quickly improve output quality and enhance the overall system response speed and resource utilization without frequently updating model parameters.

[0425] In another implementation, the server can use feedback information to retrain or fine-tune the generative AI model. The server can construct training samples from "prompt statement + expected output + feedback score" and update the model's weights using supervised fine-tuning or reinforcement learning methods based on human feedback. During training, the server can use a combination of loss functions (e.g., cross-entropy loss and rating-based reward loss) and optimize the model parameters through backpropagation and gradient updates. In this way, the server can gradually guide the model to generate outputs that are highly rated by the user, thereby technically reducing the average error and improving the suggestion matching accuracy.

[0426] VII. Technical Effects and Causal Relationships By centrally performing data preprocessing, analysis, prompt statement construction, and generative artificial intelligence model invocation, the server offers the following technical advantages compared to traditional methods that rely on simple rule judgments or template filling on the terminal side: 1. By aggregating time-series physiological indicators and handling outliers, the server reduces the amount of data that needs to be traversed for each analysis, shortens the analysis time, and improves the processing speed; 2. By providing generative AI models with prompts containing rich features and structured constraints, the server improves the relevance and stability of the output text, reduces the generation of irrelevant content and redundant sentences, thereby reducing the amount of data in the network transmission and terminal rendering stages. 3. By generating avatar display data uniformly on the server side, the terminal only needs to render and play according to the control script, which reduces the computing pressure on the terminal and is conducive to realizing complex avatar interactions on low-performance terminals; 4. By using feedback-driven prompt statement templates and adaptive updates to model parameters, the server can automatically optimize its internal processing flow without relying on manual adjustments to each rule, thereby achieving system-level accuracy improvements and reduced response times.

[0427] This system does not merely automate manual operations, but optimizes the way data is organized, models are invoked, and resources are allocated within the computer. This makes the collaborative computing between the server and the terminal more efficient and accurate in the specific scenario of health management and daily life management, thus reflecting an improvement in computer technology itself.

[0428] VIII. Other Implementation Forms and Variations In one variation, the server can employ different types of generative artificial intelligence models, such as encoder-decoder structures or instruction fine-tuning models, but still input health analysis results and schedule information into the model through a similar prompting statement construction mechanism.

[0429] In another implementation, the server can use different database types (such as columnar databases or time-series databases) to optimize the access performance of time-series physiological data, thereby further improving the concurrent processing capabilities in large-scale user scenarios.

[0430] In another embodiment, the terminal can be the wearable device itself, and the server communicates with the device through a gateway. The core of this invention lies in the centralized analysis, prompt statement construction, and generative artificial intelligence model invocation logic on the server side, which is not limited by the specific terminal form.

[0431] Through the above-described embodiments, the server, terminal, and user form a closed loop in the system of this invention: the server is responsible for high-intensity computation and model inference, the terminal is responsible for data collection and presentation, and the user drives the system to continuously optimize through natural interaction and feedback, thereby enabling the invention to be actually implemented and achieve the expected technical effects.

[0432] use Figure 13 The processing flow is explained.

[0433] Step 1: Users complete account registration and basic information entry on the terminal.

[0434] Users enter basic information such as name, age, height, weight, and past medical conditions into the terminal.

[0435] Input: Text information entered by the user in the terminal interface.

[0436] The terminal performs format validation on the input (e.g., checking if the age is an integer, and whether the height and weight are within a preset range), and temporarily stores it locally as a structured data object.

[0437] The terminal sends structured user information to the server via HTTPS.

[0438] Output: A data packet containing basic user information sent to the server.

[0439] Step 2: The server receives and stores basic user information.

[0440] The server receives user basic information requests uploaded by the terminal from the network interface.

[0441] Input: JSON data containing basic user information.

[0442] The server parses the JSON, verifies the integrity and data type of the fields, and generates an internal user record structure. The server writes a new record to the user information table and assigns a unique user identifier to the user.

[0443] Output: Response data containing the user identifier and authentication token, sent back to the terminal.

[0444] Step 3: The terminal completes the pairing of the wearable device.

[0445] The terminal scans for nearby wearable information processing devices via Bluetooth and displays a list of devices.

[0446] Input: Bluetooth identifier and service information broadcast by the wearable device.

[0447] The terminal initiates pairing based on the device selected by the user, reads the UUID of features such as heart rate service and step counting service, and generates a device binding information structure locally.

[0448] The terminal sends the device identifier and user identifier to the server via HTTPS.

[0449] Output: A message containing the device identifier and binding status is sent to the server, while the device connection parameters are saved locally on the terminal.

[0450] Step 4: The server records the device binding relationships.

[0451] The server receives a device binding request from the terminal.

[0452] Input: User ID and Device ID.

[0453] The server writes a binding record into the device information table, establishing an association between the device identifier and the user identifier.

[0454] Output: A successful binding response is sent to the terminal to notify the user that the binding is complete.

[0455] Step 5: The terminal periodically collects health-related information.

[0456] The terminal reads data such as heart rate, steps, and body temperature from the wearable device via Bluetooth at preset time intervals.

[0457] Input: Current sensor data of the wearable device (heart rate, cumulative steps, temperature).

[0458] The terminal adds a local timestamp and device identifier to each data entry and stores them as records in the local database. The terminal can perform simple deduplication or merging of consecutive data to reduce redundancy.

[0459] Output: A collection of local health data records, to be uploaded in batches later.

[0460] Step 6: Terminals upload health-related information to the server in batches.

[0461] When the preset upload conditions are met (such as time interval or record quantity threshold), the terminal reads the unuploaded health records from the local database.

[0462] Input: Multiple health data records stored locally (timestamp, heart rate, steps, body temperature, device identifier).

[0463] The terminal packages multiple records into a JSON array and sends it to the server's data upload interface via HTTPS.

[0464] Output: An upload request message containing multiple health records.

[0465] Step 7: The server receives and preprocesses health-related information.

[0466] The server receives health data uploaded by the terminal from the network interface.

[0467] Input: Multiple timestamped records of heart rate, steps, and body temperature.

[0468] The server parses the data and performs preprocessing, including: removing abnormal records with a heart rate of 0 or outside the physiologically reasonable range; sorting the data chronologically; and downsampling or merging records that are too densely packed within the same time period. The server then writes the cleaned records into a health record table.

[0469] Output: Normalized health records stored in the database, and cleaning logs.

[0470] Step 8: The server aggregates and extracts features from time-series physiological indicators.

[0471] The server reads data from the health record table by user and time window (e.g., every 5 minutes).

[0472] Input: Multiple health records (heart rate, steps, body temperature) within a specified time range.

[0473] The server processes data for each time window: calculating average heart rate, maximum heart rate, minimum heart rate, step increment, and average body temperature; and identifying periods of consecutive step increments close to zero to estimate sedentary duration. The server then writes the aggregated and characteristic results into a health summary table.

[0474] Output: Summary records of health data including time window-level statistics and features.

[0475] Step 9: The server generates health status and activity level analysis results.

[0476] The server reads aggregated features from the health summary table for a specific analysis period (e.g., the last 7 days).

[0477] Input: Features such as average heart rate over multiple days, maximum heart rate, total steps, and sedentary time.

[0478] The server uses pre-trained classification and regression algorithms (such as logistic regression or decision trees) to classify activity levels and calculates indicators such as resting heart rate trends, average daily steps, and differences from recommended values; the server combines these results into a structured analysis result object.

[0479] Output: Health analysis results data including activity level, risk warnings, trend descriptions, etc.

[0480] Step 10: The server retrieves and integrates the schedule information.

[0481] The server reads the user's medication plan, exercise plan, and other daily reminders from the database of the schedule management module.

[0482] Input: User's schedule records (event type, time, repetition rules).

[0483] The server aligns schedule information with health analysis results by time and category, for example, associating "weekly Wednesday exercise plan" with analysis tags indicating recent lack of exercise, which is then used to generate subsequent recommendations.

[0484] Output: A comprehensive contextual data structure containing health analysis results and schedule information.

[0485] Step 11: The server provides prompts for constructing generative artificial intelligence models.

[0486] The server combines basic user attributes (age, gender, height, weight), health analysis results, and schedule information into a natural language description.

[0487] Input: Basic user information, health analysis results data, and schedule information.

[0488] The server generates prompts based on predefined templates, including explanations and instructions: For example, the server generates the following exercise suggestion message: "You are a professional health management consultant. Below is a user's basic information and health analysis results for the past 7 days. Please generate an exercise suggestion for the user that is no more than 200 words based on this information. Requirements:" 1) Use plain and easy-to-understand Chinese; 2) Provide suggested daily exercise types, durations, and approximate intensity; 3) Use a gentle and encouraging tone, and avoid using threatening language.

[0489] User basic information: Male, 35 years old, height 170cm, weight 78kg.

[0490] Health analysis results: Average daily steps: 4500; prolonged sitting time; resting heart rate slightly higher than the average for peers. Output: Prompt text in natural language.

[0491] Step 12: The server invokes a generative artificial intelligence model to generate response text.

[0492] The server passes the prompt statements as a sequence of input to the generative artificial intelligence model service.

[0493] Input: Prompt text and optional structured features (such as numerical feature vectors).

[0494] Generative AI models perform vector operations within their neural network structure: embedding layers map words to vector space, multi-head self-attention layers calculate attention weights between words, feedforward layers perform non-linear transformations to progressively generate the probability distribution of the next word, and finally output complete sentences according to a sampling strategy. The server receives the natural language suggestions output by the model and performs post-processing such as length truncation and sensitive word filtering.

[0495] Output: Response text containing health management advice and / or daily life advice.

[0496] Step 13: The server generates avatar display data.

[0497] The server analyzes sentence structure and sentiment based on the response text.

[0498] Input: Post-processed response text.

[0499] The server segments the text into sentences or phrases, assigns speech parameters (speech rate, pauses) and action labels (such as nodding, smiling) to each segment; the server constructs an avatar display data structure that includes text, speech configuration, action scripts, and emotion labels.

[0500] Output: An avatar display data object used to drive the terminal avatar display.

[0501] Step 14: The server infers the user's emotional state and constructs avatar-based response prompts.

[0502] The server extracts emotion-related features from health-related information, schedule execution, and historical feedback information.

[0503] Input: Recent health characteristics, task completion rate, feedback scores, and other data.

[0504] The server uses simple models or rules (such as threshold-based rules and weighted scoring) to infer that the user may be in an emotional state such as tension, fatigue, or positivity; based on this emotional state, the server generates another type of prompting statement to guide the generative artificial intelligence model to generate avatar response content, such as words of comfort or encouragement.

[0505] Output: Emotion-related prompt text and presumed emotion label.

[0506] Step 15: The server generates avatar response content based on emotion prompts.

[0507] The server inputs sentiment-based prompts into the generative artificial intelligence model.

[0508] Input: A prompt text describing the user's emotional state and expected response style.

[0509] Generative AI models generate natural language responses that match emotional states, such as comforting words, encouraging words, or lighthearted conversational content; the server performs security checks and length limits on the responses.

[0510] Output: Emotional response text content used for avatar expression.

[0511] Step 16: The server will eventually display the data and send it to the terminal.

[0512] The server integrates the avatar display data corresponding to the response text with the emotional response content into a unified data packet.

[0513] Input: Avatar displays data objects and emotional response text.

[0514] The server returns the integrated data to the terminal via HTTPS, including text, voice, and motion control parameters.

[0515] Output: Control data for the avatar display sent to the terminal.

[0516] Step 17: The terminal renders an avatar and presents suggestions and responses to the user.

[0517] The terminal receives avatar display control data sent by the server.

[0518] Input: An avatar control object containing text, voice parameters, and action scripts.

[0519] The terminal loads the avatar model on the display screen, drives the changes in expression and posture according to the action script, and calls the text-to-speech interface to convert the text into speech output; the terminal synchronizes the speech playback with the avatar's lip-sync animation, so that the user can receive suggestions and emotional responses through both sight and hearing.

[0520] Output: The virtual avatar and its voice content displayed on the terminal screen and speakers.

[0521] Step 18: Users can view or listen to suggestions and provide feedback.

[0522] Users can read text displayed in the avatar or listen to audio on the terminal.

[0523] Input: Suggested content and emotional response displayed on the terminal.

[0524] Users can rate the quality of the current suggestion by clicking the "Helpful / Not Helpful" button, selecting a star rating, or entering a comment; the terminal converts the user's feedback into structured feedback data.

[0525] Output: Records of feedback information generated locally on the terminal.

[0526] Step 19: The terminal uploads feedback information to the server.

[0527] After the user submits feedback, the terminal sends a feedback request to the server via HTTPS.

[0528] Input: Feedback data including user IDs, suggestion IDs, ratings, and text comments.

[0529] After the terminal successfully sends the message, it can mark the feedback record as synchronized locally.

[0530] Output: The feedback data message sent to the server.

[0531] Step 20: The server stores feedback information and uses it for prompts and model optimization.

[0532] The server receives and parses the feedback data uploaded by the terminal.

[0533] Input: Structured feedback information.

[0534] The server writes feedback into a feedback information table and categorizes it according to suggestion type, prompt statement template, and model version. The server periodically performs statistical analysis on the feedback, calculates the average score and adoption rate of different templates, and automatically adjusts the content of prompt statement templates or model running parameters based on the statistical results, thereby changing prompt statement constraints or output strategies in subsequent calls.

[0535] Output: Updated suggestion template parameters and / or generative AI model configuration for the next round of suggestion generation.

[0536] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0537] With the development of human-computer interaction and artificial intelligence technologies, systems that utilize virtual display objects to communicate with users are gradually increasing. However, existing technologies have the following problems: First, the generation of traditional virtual characters or avatars is mostly based on fixed templates or simple parameter configurations. It cannot accurately parse the complex requirements for appearance and behavioral characteristics from the prompts entered by users in natural language, resulting in limited personalization of virtual display objects and a poor user experience.

[0538] Second, existing systems typically use only single-modal data (such as only facial images or only voice tone) for emotion recognition in terms of emotional interaction. They cannot integrate multi-source state information such as physiological sensor data, image data, and voice data, which makes it difficult to accurately grasp the user's true emotional and physical state and provide appropriate feedback in a timely and effective manner.

[0539] Third, existing dialogue systems mostly use fixed scripts or simple rules to generate response content based on limited contextual information. They lack a mechanism to transform user status summary information into prompt statements and drive generative artificial intelligence models to dynamically generate response content, making it difficult for the system response to achieve fine-grained adaptive adjustment for different user states.

[0540] Fourth, existing systems generally lack the ability to self-optimize interaction strategies based on users' historical operations and behavioral patterns. They cannot automatically adjust the conditions for generating subsequent prompts and the response methods of virtual display objects based on users' acceptance of past responses, thus failing to continuously improve the human-computer interaction effect and the efficiency of system resource utilization during long-term interaction.

[0541] Therefore, how to achieve the following within the same system: (1) Accurately parse user prompts using natural language processing and generate feature data for controlling the visual attributes and behavioral characteristics of virtual display objects; (2) Use multi-source state information to accurately estimate the user's emotional state and physical state, and abstract it into state summary information that can be effectively used by generative artificial intelligence models; (3) Based on the above status summary information, prompt statements are automatically constructed to drive the generative artificial intelligence model to dynamically generate response content. (4) Based on user operation information and behavior history, iteratively optimize the conditions for generating prompt statements and the response methods. Improving the adaptability, robustness, and resource utilization efficiency of virtual interactive systems at both the algorithm and system architecture levels has become an urgent issue for computer technology in this field.

[0542] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.

[0543] In this invention, the server includes a processing unit for receiving natural language prompts from the user and parsing the natural language prompts to generate control information for determining the visual attributes and behavioral characteristics of a virtual display object; a processing unit for generating feature data representing the constituent elements of the virtual display object based on the control information and sending the feature data to an information processing terminal; a processing unit for acquiring state information including physiological sensor data, image data, and voice data and inferring the user's emotional and physical states based on the state information; a processing unit for generating prompts input to a generative artificial intelligence model using state summary information about the emotional and physical states and generating virtual display object response content based on the response results obtained from the generative artificial intelligence model; and a processing unit for acquiring user operation information and behavioral history information regarding the response content and updating the prompt generation conditions and virtual display object response methods based on the information. This can be achieved internally within a computer using technical means: on the one hand, by jointly processing natural language prompts and multi-source state information, the appearance and behavior of virtual display objects can be precisely controlled, and the accuracy of emotion recognition can be improved; on the other hand, by constructing prompts based on state summary information and driving a generative artificial intelligence model to dynamically generate response content, while continuously adjusting the generation strategy using user behavior feedback, the human-computer interaction process can be optimized at the algorithm level, improving the system's adaptability to different users and different states, and improving the overall utilization efficiency of computing resources and the interactive experience.

[0544] A "system" refers to a computer-implemented overall technical solution consisting of one or more processing devices, storage devices, and terminal devices that communicate with it, used to perform data acquisition, analysis, and output control.

[0545] "Processing device" refers to a computer unit that includes a processor and the programs running on it, used to parse, calculate, judge, and generate control information or output results from input data.

[0546] "Information processing terminal" refers to an electronic device that communicates with a server to collect user data and present virtual display objects and related information to the user, including but not limited to smartphones, tablets, head-mounted displays or other computing devices.

[0547] "User" refers to the entity that interacts with the system, provides input data, and receives output from the system.

[0548] "Natural language prompts" refer to textual information entered by users in natural language to express needs, preferences, states, or instructions, and are used as input for generative artificial intelligence models and parsing modules.

[0549] "Parsing" refers to the process of using algorithms to perform syntactic analysis, semantic understanding, feature extraction, and structuring of input data in order to obtain information that can be used for subsequent calculations and control.

[0550] "Control information" refers to a set of data generated based on the analysis results, used to determine the internal parameters of virtual display objects, such as their visual attributes, behavioral characteristics, or response methods.

[0551] "Virtual display objects" refer to digital characters or graphic entities presented on a display interface for interaction with users, including their appearance, expressions, and actions.

[0552] "Visual attributes" refer to the appearance characteristics of virtual display objects on the display interface, including parameters related to visual performance such as shape, color, material, clothing, and hairstyle.

[0553] "Behavioral characteristics" refer to the dynamic features of virtual display objects during interaction with users, such as action patterns, postures, interaction methods, and dialogue styles.

[0554] "Constituent elements" refer to the various components that make up a virtual display object, including but not limited to the head, body, limbs, clothing, accessories and their corresponding attribute parameters.

[0555] "Feature data" refers to a set of structured data or parameters used to describe the constituent elements, visual attributes, and behavioral characteristics of virtual display objects, which can be directly used by the rendering engine or display module.

[0556] "Status information" refers to a set of data related to a user's current physiology, emotions, and environment, including physiological sensor data, image data, voice data, and other behavioral data.

[0557] "Physiological sensor data" refers to data related to a user's physical condition acquired by physiological detection devices, including but not limited to heart rate, activity level, sleep data, or other vital signs data.

[0558] "Image data" refers to static image or video frame data that contains information about a user's face or posture, collected by an imaging device.

[0559] “Voice data” refers to audio signals acquired by a sound acquisition device that contain the content and intonation of a user’s speech.

[0560] "Emotional state" refers to the psychological or emotional state of a user at a specific time, inferred from state information, such as joy, sadness, anger, tension, relaxation, etc.

[0561] "Body condition" refers to the user's physiological or health status inferred from physiological sensor data and related information, such as fatigue, excitement, or adequate rest.

[0562] "State summary information" refers to an abstract or compressed representation of information generated based on emotional and physical states, used to briefly describe a user's current overall state.

[0563] "Generative AI models" refer to AI models that can automatically generate text, parameters, or other content based on input prompts, including deep learning-based text generation models or multimodal generation models.

[0564] "Prompt statements" refer to natural language instructions or descriptive text input into a generative artificial intelligence model to constrain the generated content. These can include information such as scene descriptions, task requirements, and style limitations.

[0565] "Response result" refers to the text content, structured data, or other generated content that a generative artificial intelligence model outputs after receiving a prompt statement, which is used to drive the behavior of virtual display objects.

[0566] "Response content" refers to the specific dialogue text, notification information, or explanatory information generated based on the response result and combined with system logic, which is presented to the user by the virtual display object.

[0567] "Facial expression information" refers to a set of parameters used to control the facial expressions and emotional expressions of virtual display objects, including expression types and their intensity such as smiling, frowning, and surprise.

[0568] "Motion information" refers to a set of parameters used to control the body movements, posture changes, or interactive behaviors of virtual display objects, including action type, duration, and triggering conditions.

[0569] "Display behavior" refers to the overall performance of virtual display objects on the terminal display interface, including appearance, facial expression changes, action playback, and interaction with user interface elements.

[0570] "Operation information" refers to the interactive data generated when a user responds to the system output through a terminal, including records of clicks, selections, inputs, voice commands, or other interactive operations.

[0571] "Behavioral history information" refers to the user interaction trajectory and usage patterns recorded by the system over a period of time, including statistical or time-series data such as acceptance of response content, dwell time, and interaction frequency.

[0572] "Prompt statement generation conditions" refers to the rules, parameters, and context information used when generating prompt statements, including user preferences, historical feedback, current status summary information, and system policy configuration.

[0573] "Response style" refers to the form and style of expression used by virtual display objects when outputting information to users, including language style, information length, presentation rhythm, tone strength, and whether it is accompanied by actions or expressions.

[0574] In the following description, the server performs the main data processing, model inference, and control logic, the terminal performs data acquisition and display control, and the user interacts with the system naturally. The embodiments of this invention revolve around the technical solution of "how the server uses a generative artificial intelligence model and prompt statements to control virtual display objects, and adaptively adjusts them in conjunction with multi-source state information."

[0575] I. System Overall Structure A server includes one or more general-purpose processors, a graphics processor, storage devices, and a network communication interface. The server runs an operating system and several application modules on the processor, including: a natural language processing module, a state estimation module, a generative artificial intelligence invocation module, a feature data generation module, a behavior decision-making module, and a user preference learning module, etc.

[0576] The terminal includes a display device, an input device, a camera device, a sound acquisition device, and an optional physiological sensor interface. The terminal runs a local visual presentation program, such as a virtual object display module based on a 3D rendering engine (e.g., general-purpose 3D engine software), and a client module that communicates with the server.

[0577] Users input prompts, voice messages, and other operational information through the terminal, and view the appearance and response content of the virtual display objects through the terminal's display interface.

[0578] II. Program Modules and Data Structures The server maintains multiple data structures in the storage device to improve processing efficiency and achieve technical effects: 1. The server stores the attribute vector of each user in the "user configuration database". The attribute vector includes the user's basic attribute fields (such as the encoding of age group and gender category), preference tag vector (such as the preference weight of "gentle" and "lively" personality), and historical emotion distribution statistics (such as the proportion of "joy", "sadness" and "stress" in a recent period).

[0579] 2. The server organizes the feature data into a multi-layer parameter structure in the "Virtual Display Object Configuration Table", which includes visual attribute parameters (color, material, geometric shape parameters), behavioral feature parameters (action library index, behavioral style weight) and expression parameters (expression type One-Hot encoding, amplitude parameter, duration parameter).

[0580] 3. The server stores physiological sensor data, image feature vectors, and speech feature vectors in a time-series format in the "State Information Cache." Each record includes at least a timestamp field, a user identifier field, and a feature vector field for the corresponding modality. The server uses this structure to quickly retrieve data by time window during subsequent state estimation, thereby reducing database query overhead and improving processing speed.

[0581] 4. The server maintains several natural language templates in the "Prompt Statement Template Library" to construct prompt statements input to the generative AI model. When combining state summary information and user preference information, the server can select appropriate templates and fill in the specific content, thereby reducing the burden on the generative AI model while maintaining structural stability and improving the stability and consistency of generation.

[0582] III. Natural Language Parsing and Control Information Generation The server runs a deep learning-based sequence labeling model and dependency parsing algorithm in its natural language processing module. The server converts the user-input natural language prompts into word sequences and maps each word to a vector representation through an embedding layer. The server uses a bidirectional recurrent neural network or a self-attention structure to encode the word sequences, obtaining context-sensitive word representation vectors. The server then labels each word with a semantic category, such as "color," "personality adjective," or "entity type," through a classification layer or a conditional random field layer.

[0583] Based on the above annotation results, the server constructs an intermediate representation containing attribute control information in the form of key-value pairs, such as "hair_color:blue" and "personality:gentle". In this process, the server uses a strategy that combines rule sets and statistical models: on the one hand, it uses a mapping table to map color words to standard color codes, and on the other hand, it uses clustering results to map personality adjectives to predefined behavioral feature dimensions (e.g., "gentle" is mapped to multi-dimensional parameters such as "low aggression, slow speech, and soft expression").

[0584] The server uses this structured parsing process to map natural language prompts into computable control information, thereby avoiding traditional configuration methods based on fixed menus or forms, increasing the degree of freedom of expression, and ensuring the stability of parsing through model structure and rule constraints.

[0585] IV. Feature Data Generation and Terminal Display Control In the feature data generation module, the server maps control information to the internal parameter space of the virtual display object. The server uses parameter tables to map logical attributes (such as "blue hair" and "gentle personality") to rendering parameters and behavioral parameters. For example, the server maps "blue hair" to a specific color vector and hairstyle model index, and "gentle" to a combination of behavioral features with low action weights and an expression that leans towards smiling.

[0586] The server packages the generated feature data into structured data records and sends them to the terminal. The terminal reads these parameters in the rendering engine, loads the corresponding model resources into memory, and adjusts materials, skeletal animation weights, etc., according to specified parameters. Based on the unified parameter structure provided by the server, the terminal can efficiently render virtual display objects locally, avoiding the execution of complex parsing logic on the terminal, thereby reducing the terminal's computational burden and improving rendering stability.

[0587] V. Multi-source state information processing and sentiment inference In the state estimation module, the server performs feature extraction and normalization on physiological sensor data, image data, and voice data respectively, and then concatenates or aligns these features into a state vector on a uniform time scale.

[0588] In terms of physiological data processing, the server can perform filtering, resampling, and variability analysis on the heart rate sequence using standard signal processing steps. The server calculates short-term and long-term averages, the standard deviation of heart rate variability, and activity levels, combining these statistical results into a physiological feature vector. This processing method reduces the impact of transient noise on judgment and provides more stable input for subsequent models.

[0589] In image processing, the server extracts facial features using a convolutional neural network (CNN). This CNN comprises multiple layers of convolutional, pooling, and normalization layers to extract expression-sensitive features from the original image. At the network's end, a fully connected layer outputs a multidimensional emotion distribution, representing probabilities such as "joy," "sadness," "anger," and "neutrality." During the training phase of this neural network, the server uses a cross-entropy loss function, taking the difference between the predicted distribution and the manually labeled data as the loss, and updates the network weights through backpropagation. The server can also employ data augmentation strategies, such as randomly rotating or varying the brightness of training samples, to improve the network's robustness to environmental changes.

[0590] In terms of speech processing, the server converts the speech signal into a spectrogram or Mel-frequency cepstral coefficient feature, and extracts speech emotion features through a one-dimensional or two-dimensional convolutional network. The server also uses supervised learning methods to train the network to output the probabilities of different emotion categories.

[0591] In the multimodal fusion stage, the server performs weighted fusion of physiological feature vectors, image emotion distribution, and speech emotion distribution. The server can employ a simple linear weighting strategy or a multilayer perceptron model, cascading the feature vectors of each modality as input and outputting a unified emotion state vector and body state estimate through nonlinear transformation. This design introduces learnable weights, enabling the system to automatically adjust the contribution ratio of different modalities based on historical performance, thereby improving the accuracy and robustness of emotion recognition within the network's internal structure.

[0592] VI. Combining State Summary Information with Generative Artificial Intelligence Models After obtaining the emotional and physical states, the server encodes this state information into a compact summary. For example, the server compresses "average heart rate over the past 30 minutes, sedentary percentage, main emotional category, and confidence level" into a short text description string, while internally retaining a structured numerical representation for algorithmic decision-making.

[0593] When generating a prompt message, the server uses a prompt message template plus variable population. For example, when the server detects that a user is under high stress and has been sitting for too long, the server can construct the following prompt message: "You are a virtual assistant that cares about users' health."

[0594] User's current status: average heart rate of 95 over the past 30 minutes, 90 minutes of continuous sitting, mood is 'stressed'.

[0595] Task: Please use a brief, gentle tone to remind the user to rest for 5 minutes and drink water. You can suggest one or two simple stretching exercises. Keep the total word count under 40 words. In this way, the server transforms computable states into natural language prompts that can be understood and utilized by generative AI models. Because the server rewrites state data into a structured natural language context, generative AI models can reason and generate in a richer semantic space. Compared to simple rule systems, they can produce a wide variety of outputs that still meet constraints based on different combinations of states, thereby improving the adaptability of interactive content.

[0596] When the server recognizes that a user's emotional state is "sad / depressed", the server can construct another prompt message: "You are a gentle virtual companion who needs to comfort users who are feeling down."

[0597] The user just said: 'Today wasn't going well.' Task: Respond to users with sincerity and understanding, acknowledge their feelings, offer appropriate encouragement, but avoid forcing positive responses. Keep the response to 50 words or less. The server controls the output style and length of the generative AI model through this mechanism of combining templates and variables, reducing the risk of generating irrelevant or excessively long content. At the same time, it enables the model to adopt different generation strategies for different scenarios, reflecting the engineering constraints and optimizations on the generative reasoning process.

[0598] VII. Structure and Learning Methods of Generative Artificial Intelligence Models In the generative AI model portion, the server can employ a sequence-to-sequence model structure based on a self-attention mechanism. During the training phase, the server pre-trains the model using a large-scale text corpus, enabling the model to learn common language patterns, semantic associations, and pragmatic rules. During pre-training, the server constructs a training task by masking parts of the input words and predicting the masked words. It uses a cross-entropy loss function to measure the difference between the predicted and true distributions and updates the network parameters using a gradient descent algorithm.

[0599] The server can fine-tune the pre-trained model in specific scenarios. The server uses dialogue samples, health reminder texts, and emotional reassurance statements from the target application scenario as fine-tuning data to make the model more suitable for the application scenario of this invention. During the fine-tuning process, the server can add additional control signals, such as "tone labels" and "length constraints," to further improve the model's compliance with the constraints in the prompt statements.

[0600] Through the above design, the server adapts the parameter distribution within the generative artificial intelligence model to the various output styles required by this invention. Combined with the structured construction of prompt statements, it achieves controllability and efficiency in the natural language generation process.

[0601] 8. Response content generation and facial expression / action mapping After receiving the response results from the generative artificial intelligence model, the server maps the text content to the facial expressions and action parameters of the virtual display objects through the behavior decision module. The server maintains a "semantic-to-behavior mapping table" in storage, which associates typical semantic tags (such as "comfort", "encouragement", "remind to rest", "recommend products") with a set of action patterns, such as "gentle smile + slow nod" and "excited wave + wide eyes".

[0602] The server performs lightweight semantic classification on the generated text, such as through keyword matching or a small classification network, to determine which semantic function the text belongs to, and then selects the corresponding behavior pattern according to a mapping table. In some implementations, the server can add weights to the behavior pattern based on emotional state, for example, reducing the amplitude of movements when the user is "very tired" to avoid causing visual fatigue.

[0603] The server encapsulates the final text content, facial expression parameters, and action parameters together and sends them to the terminal. The terminal, in its rendering engine, uses these parameters to drive the virtual display object to play the corresponding animation and display the text, thus creating a consistent and coherent multimodal interactive experience.

[0604] IX. User Feedback and Adaptive Optimization In its user preference learning module, the server continuously records user actions and behavioral history related to responses. For example, the server records whether the user viewed the full prompt, quickly closed the dialog window, performed a recommended action, or asked further questions. Based on this data, the server updates the prompt generation conditions and response methods.

[0605] Servers can use gradient boosting trees, linear models, or reinforcement learning algorithms to learn which prompt template, tone label, or length constraint to choose in different states to achieve higher user response rates. By maintaining policy parameters internally and referring to these parameters each time a prompt is generated, the server gradually adjusts its system behavior to better suit individual user preferences.

[0606] This feedback-driven adaptive optimization mechanism not only improves the user experience but also technically enables dynamic adjustment of generative AI model invocation strategies and optimization of resource utilization. For example, when the server detects that users generally skip overly long texts, it will generally use shorter length constraints in subsequent prompts, thereby reducing the length of generated content, lowering network transmission and rendering burdens, and improving the overall system response speed.

[0607] 10. Technological Effects and Improvements in Computer Technology Through the specific structure and algorithm described above, the server not only automates the virtual interaction process but also brings improvements at the following technical levels: 1. The server utilizes a multimodal fusion-based sentiment and state estimation network, which improves sentiment recognition accuracy and reduces false positives caused by single-channel noise compared to traditional systems using only a single modality. This directly improves the relevance of subsequently generated content.

[0608] 2. By compressing status information into a status summary and using it in the prompt statement template, the server makes the data used during transmission and generation more compact, reducing network communication load and the processing burden of generative artificial intelligence models, thereby improving the overall response speed.

[0609] 3. By using predefined mapping structures and structured feature data, the server achieves efficient conversion from natural language attribute descriptions to rendering parameters, avoiding repeated parsing of complex natural language on the terminal, improving rendering efficiency on the terminal side, and reducing computational requirements.

[0610] 4. By introducing user feedback-driven strategy learning, the server continuously adjusts the conditions for generating prompts and the response methods, making the calls to the generative artificial intelligence model more targeted. This reduces redundant calls and improves the efficiency of computing resource utilization in long-term operation.

[0611] 5. The server employs specific network structures, loss functions, and training strategies during model training and inference, thereby optimizing the model's internal parameters and improving its generalization ability. This technically enhances the computer system's ability to process natural language and multimodal state information.

[0612] Through the above-described embodiments, this invention organically combines generative artificial intelligence models with prompt statements, and through a clear data structure and algorithm flow, it translates abstract emotional interaction needs into specific internal computer processing steps and terminal device control behaviors, thereby realizing a technical solution that can be directly deployed on real devices, rather than just an abstract description of business processes.

[0613] use Figure 14 The processing flow is explained.

[0614] Step 1: Users input natural language prompts through the terminal and generate raw interactive data.

[0615] Input: User's natural language text (prompt statements), voice data, image data, and optional raw data from physiological sensors.

[0616] Specific actions: The user types prompts such as "Hair is blue, personality should be gentle" or "A little tired today" into the text input box of the terminal, or speaks into the microphone of the terminal and naturally expresses facial expressions in front of the terminal's camera; the terminal calls the keyboard interface, microphone interface, camera interface, and connection interface with wearable devices to collect the above text, audio streams, image frames, and physiological signals such as heart rate and activity level, and packages this data into a request message with timestamps and user identifiers.

[0617] Output: A raw interactive data packet containing user prompts, audio clips, image frames, and physiological data, ready to be sent to the server.

[0618] Step 2: The terminal uploads the raw interactive data to the server and completes preliminary preprocessing.

[0619] Input: The raw interactive data packet generated in step 1.

[0620] Specific actions: The terminal sends data packets to the server's predetermined interface through a network communication module (such as an HTTP or WebSocket protocol stack); before sending, the terminal performs basic compression on audio and images, such as encoding audio into a compressed format and converting images to JPEG format to reduce data size; the terminal performs simple sampling or packaging of physiological data and aggregates data points at fixed time intervals to reduce redundancy.

[0621] Data processing / data calculation: The terminal performs time aggregation on physiological data based on the sampling period and performs encoding and compression operations on audio and video data to generate media data with small volume but sufficient information integrity.

[0622] Output: The compressed and packaged interactive data request is sent to the server over the network, and the server receives the data stream in the interface module.

[0623] Step 3: The server parses user prompts and generates control information.

[0624] Input: Natural language prompt text from the terminal.

[0625] Specific actions: In the natural language processing module, the server segments the prompt statement into a word sequence, uses a word embedding model to convert each word into a vector representation, and then performs forward computation on the vector sequence through a deep neural network (such as a sequence model with a self-attention mechanism); the server performs label prediction on the network output to identify color words, personality adjectives, object types, etc.; the server further applies a rule mapping table to map words such as "blue" to color codes and "gentle" to personality feature dimensions.

[0626] Data processing / data computation: The server performs syntactic analysis and semantic annotation on the prompt statements, calculates word vectors, attention weights and classification probabilities, and converts unstructured text data into structured control information, such as key-value pairs "hair_color=blue" and "personality=gentle".

[0627] Output: An intermediate control information data structure representing the visual attributes and behavioral characteristics of virtual display objects.

[0628] Step 4: The server generates virtual display object feature data based on the control information and sends it to the terminal.

[0629] Input: The control information data structure output from step 3.

[0630] Specific actions: The server looks up the parameter mapping table in the feature data generation module and maps logical attributes to rendering and behavior parameters. For example, "hair_color=blue" is mapped to specific RGB color values ​​and hairstyle IDs, and "personality=gentle" is mapped to motion amplitude parameters, facial expression intensity parameters, and speech rate parameters. The server combines different categories of parameters into a unified feature vector or configuration object.

[0631] Data processing / data computation: The server performs table lookup operations and vector assembly, encodes high-level semantic attributes into low-level numerical parameters, and constructs the feature data structure of virtual display objects; the server normalizes the parameters so that the terminal rendering engine can use them directly.

[0632] Output: Virtual display object feature data containing information such as geometry, material color, facial baseline, and behavioral style, which is sent to the terminal via the network.

[0633] Step 5: The terminal renders and displays virtual display objects based on feature data.

[0634] Input: Virtual display object feature data sent by the server.

[0635] Specific actions: The terminal loads the specified model resources in the 3D rendering engine, sets visual parameters such as material color and hairstyle according to feature data, and applies initial expressions and postures; the terminal draws the rendering results on the display screen and presents the updated virtual display object to the user.

[0636] Data processing / data calculation: The terminal performs vertex transformation, rasterization, shading calculation and other operations in the graphics rendering pipeline, converting numerical parameters and model data into screen pixels; the terminal also adjusts the animation playback rate and loop mode according to the behavior style parameters.

[0637] Output: Personalized virtual display objects visible in the terminal display interface, laying the visual foundation for subsequent interactions.

[0638] Step 6: The server extracts features from multi-source state information and infers emotions and physical states.

[0639] Input: Physiological sensor data, image data, and voice data uploaded from the terminal.

[0640] Specific actions: The server performs filtering and statistical operations on physiological data to calculate features such as average heart rate, heart rate variability, and activity level; the server inputs images into a convolutional neural network to extract high-dimensional facial expression features and calculates the probability distribution of each emotion category through the output layer; the server extracts the spectrum or Mel-frequency cepstral coefficients from the speech data and inputs them into the speech emotion network to obtain the speech emotion distribution.

[0641] Data processing / data computation: The server performs feature extraction operations (convolution, pooling, normalization, etc.) and classification operations (fully connected, softmax) on each modality. Then, it concatenates or weights and fuses the multimodal feature vectors to calculate a comprehensive emotional state vector and physical state index, such as "emotion = stress, confidence 0.8; physical state = fatigue".

[0642] Output: State estimation results including emotional state labels, physical state labels, and their confidence scores.

[0643] Step 7: The server generates a status summary and constructs prompts for inputting the generative artificial intelligence model.

[0644] Input: The emotional and physical state estimation results output from step 6, as well as the user's basic attributes and preference information.

[0645] Specific actions: The server converts the state estimation results into a concise text description or structured description, such as "average heart rate of 95 over the past 30 minutes, 90 minutes of continuous sitting, and stress mood"; the server selects an appropriate prompt statement template, fills the state description into the corresponding position in the template, and generates a complete prompt statement.

[0646] Data processing / data computation: The server maps numerical state data into semantic text fragments and performs template filling operations to combine these fragments into prompts with context, task constraints and style descriptions, which are used to drive generative artificial intelligence models.

[0647] Output: Natural language prompts containing specific task constraints and state descriptions, for example: "You are a virtual assistant that cares about users' health."

[0648] User's current status: average heart rate of 95 over the past 30 minutes, 90 minutes of continuous sitting, mood is 'stressed'.

[0649] Task: Please use a brief, gentle tone to remind the user to rest for 5 minutes and drink water. You can suggest one or two simple stretching exercises. Keep the total word count under 40 words. Step 8: The server invokes a generative artificial intelligence model to generate response text based on the prompts.

[0650] Input: The prompt statement generated in step 7.

[0651] Specific actions: The server encodes the prompt statement as an input sequence into a vector representation, and performs forward inference using a pre-trained and fine-tuned generative artificial intelligence model (a language model based on a self-attention structure); the server calculates the contextual representation of each position through an attention mechanism within the model, and generates text word by word through the output layer; the server determines the end of generation based on length constraints and stop markers.

[0652] Data processing / data computation: The server performs encoding operations and self-attention matrix multiplication operations on the input prompt statements, samples or selects the words with the highest probability according to the probability distribution, and gradually constructs the output sequence to obtain the response text that meets the task requirements, such as comforting statements, health reminders or product recommendations.

[0653] Output: A natural language response text, such as "You've been sitting for a long time. Stand up and walk around for a few minutes, and have some water to relax your neck and shoulders." Step 9: The server maps the response text to the facial expressions and action parameters of the virtual display object and issues control commands.

[0654] Input: The response text generated in step 8 and the current emotional state information.

[0655] Specific actions: The server performs a brief semantic classification of the response text to determine its functional category (e.g., "comfort", "reminder", "recommendation"), selects the corresponding expression and action pattern from the mapping table, such as "gentle smile + slow nod"; the server fine-tunes parameters such as the amplitude and duration of the action according to the emotional state, and combines them into a control instruction package.

[0656] Data processing / data computation: The server performs keyword matching or lightweight classification operations to map text into behavior category codes, and looks up the corresponding facial expression parameters and action sequence parameters in the behavior parameter library to form a complete set of multimodal control parameters.

[0657] Output: Control commands containing response text, facial expression parameters, and action parameters are sent to the terminal via the network to drive the next behavior of the virtual display object.

[0658] Step 10: The terminal displays the response content and provides feedback on user operation information based on the control commands.

[0659] Input: Control commands issued by the server (response text + emoticon parameters + action parameters).

[0660] Specific actions: The terminal displays the response text in the display interface as a speech bubble or text area. At the same time, the rendering engine applies facial expression parameters and action parameters to drive the virtual display object to perform corresponding animations, such as smiling, nodding, or pointing. The terminal can call the text-to-speech module to convert the response text into speech for playback. The terminal displays interactive buttons, such as "I'm resting" and "Will remind you later," and captures the user's clicks or voice feedback.

[0661] Data processing / data computation: The terminal performs graphics animation interpolation calculations and audio playback control, converting abstract facial expressions and action parameters into continuous frame animations and sound signals; the terminal encapsulates user operation events into operation information records, along with timestamps and current session identifiers, providing data for subsequent server learning and optimization.

[0662] Output: The response of the virtual display objects visible and audible to the user on the terminal, as well as the interaction log containing user operation feedback. The interaction log will be sent back to the server for subsequent adaptive adjustments.

[0663] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0664] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0665] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.

[0666] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0667] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.

[0668] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.

[0669] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.

[0670] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0671] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.

[0672] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0673] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0674] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0675] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0676] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0677] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0678] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.

[0679] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0680] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0681] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0682] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0683] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0684] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0685] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0686] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.

[0687] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0688] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.

[0689] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.

[0690] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.

[0691] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0692] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.

[0693] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0694] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0695] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0696] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0697] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0698] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0699] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.

[0700] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".

[0701] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0702] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0703] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0704] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0705] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0706] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0707] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.

[0708] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12 to analyze the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12 to generate a menu using a generation AI. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12 to provide the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0709] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.

[0710] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.

[0711] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.

[0712] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0713] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.

[0714] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0715] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by a perspective equivalent to the field of vision of an average healthy person).

[0716] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0717] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0718] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0719] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0720] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0721] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.

[0722] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".

[0723] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0724] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0725] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0726] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0727] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0728] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0729] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.

[0730] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0731] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.

[0732] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see...). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.

[0733] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.

[0734] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.

[0735] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).

[0736] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.

[0737] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."

[0738] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values ​​representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.

[0739] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).

[0740] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.

[0741] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0742] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.

[0743] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.

[0744] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.

[0745] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.

[0746] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.

[0747] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.

[0748] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.

[0749] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.

[0750] In addition, the following notes are provided in response to the above explanation.

[0751] Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for parsing a prompt statement containing natural language input information obtained by a user-operated information terminal by a processing unit in an information processing apparatus, and instructing a processing unit to extract preference information and demand conditions from the prompt statement. An apparatus for the processing unit to generate prompt statements for inputting into a generative artificial intelligence model based on the extracted preference information and demand conditions, and to instruct the generative artificial intelligence model to generate feature data representing the appearance attributes, personality attributes, and behavioral attributes of a virtual object. An apparatus for sending information for visual representation contained in the generated feature data to the information terminal by the processing unit, and instructing the virtual object to be visually displayed on the information terminal; The device is used to infer the user's emotional state by the processing unit based on the user's dialogue input and physiological information or operation history continuously obtained from the information terminal, and to construct a response generation prompt statement that reflects the inferred emotional state and the personality attributes of the virtual object, and to instruct the generation of response data representing the response content and facial expressions or actions of the virtual object using the generative artificial intelligence model. A device for sending the reaction data from the processing unit to the information terminal and instructing it to be presented on the information terminal in the form of the virtual object's speech display and changes in facial expressions or actions; And an apparatus for generating regeneration prompts by the processing unit based on additional prompts from the user to update the appearance and personality attributes of the virtual object, and instructing the feature data to be updated using the generative artificial intelligence model for continuous customization of the virtual object.

[0752] (Note 2) According to the information processing system described in Appendix 1, the processing unit is configured to use natural language processing technology to extract semantic information related to preferred objects, personality tendencies, and dialogue styles from the free descriptive text contained in the input information, and to provide the semantic information as a numerical attribute vector to the generative artificial intelligence model to generate the feature data.

[0753] (Note 3) According to the information processing system described in Appendix 1, the processing unit is configured to record the time changes of the user's emotional state and the history of interaction with the user, and dynamically change the emotion regulation parameters contained in the prompt statement for response generation based on the records, thereby generating the response content, expression or action of the virtual object for the purpose of alleviating the user's feelings of isolation or anxiety.

[0754] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for acquiring user behavior and preference information in an information processing device, and determining attribute information representing user interests through machine learning processing, thereby generating user profile information; A device for generating prompt statements based on the user profile information to specify the appearance and personality information of a virtual character, and for running a generative artificial intelligence model with the prompt statements as input to generate feature information containing the appearance and personality information of the virtual character; An apparatus for generating image information or stereoscopic display information representing the virtual character based on the feature information, and sending the image information or stereoscopic display information to a terminal device for visualization in virtual space; A device for acquiring user operation information and dialogue information generated in a virtual store space, determining products that match user interests based on the user profile information and product information, and generating recommendation information presented through the virtual character; An apparatus for generating a prompt statement based on the user's dialogue information and the user profile information to specify the response content output by the virtual character, and for running a generative artificial intelligence model with the prompt statement as input to generate dialogue information containing the response content, and for controlling the voice and behavior of the virtual character according to the dialogue information. An apparatus for storing the interaction history information between the user and the virtual character, and updating the user profile information and the recommendation information based on the interaction history information, thereby continuously optimizing the appearance information, personality information and product recommendation behavior of the virtual character.

[0755] (Note 2) The information processing system according to Appendix 1 is characterized in that, The information processing device further includes: a device for parsing natural language instruction text received from the user using natural language processing technology, and reflecting the instruction text in the prompt statement, thereby changing the appearance and personality information of the virtual character.

[0756] (Note 3) The information processing system according to Appendix 1 is characterized in that, The information processing device further includes: a device for inferring the user's emotional state based on the user's dialogue content and behavioral information, generating the prompt statement according to the emotional state to run the generative artificial intelligence model, thereby adjusting the virtual character's response content and behavior, and reducing the user's psychological burden through continuous interaction with the user.

[0757] Example 2 (Note 1) An information processing system, characterized in that it comprises: A device for instructing a terminal to acquire health-related information, including user physiological information, and schedule information, and for receiving the health-related information and schedule information from the terminal; An apparatus for processing data based on the health-related information to generate analysis results about the user's health status and lifestyle habits; An apparatus for constructing prompt statements as input to a generative artificial intelligence model based on the analysis results and the schedule information, and inputting the prompt statements and the analysis results into the generative artificial intelligence model so that the generative artificial intelligence model generates response text containing health management suggestions and daily life suggestions for different users; A device for generating avatar display data as avatar voice and action content from the response text, and sending the avatar display data to the terminal so that the terminal can display the avatar in a visual manner; An apparatus for updating the prompt statements input to the generative artificial intelligence model and / or the operating parameters of the generative artificial intelligence model based on feedback information obtained from the user via the terminal, thereby gradually improving the content of the health management recommendations and the daily life recommendations; An apparatus for inferring a user's emotional state based on the health-related information, the schedule information, and the feedback information, and for generating prompt statements that correspond to the emotional state using the generative artificial intelligence model, and sending the avatar response content to the terminal.

[0758] (Note 2) The information processing system according to Appendix 1 is characterized in that, The health-related information includes time-series physiological indicators such as heart rate, steps, and body temperature obtained by wearable information processing devices. The data processing includes statistical calculation of the time-series physiological indicators, outlier detection, and activity level classification.

[0759] (Note 3) The information processing system according to Appendix 1 is characterized in that, The prompts include: explanatory text describing the user's basic attributes in natural language, the analysis results, and the schedule information, as well as instruction text specifying the word range, expression style, and number of specific action items for the generated health management suggestions and daily life suggestions, thereby constraining the content input to the generative artificial intelligence model.

[0760] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: A processing device for receiving natural language prompts from a user and parsing the natural language prompts to generate control information for determining the visual attributes and behavioral characteristics of a virtual display object; A processing device for generating feature data representing the constituent elements of the virtual display object based on the control information, and sending the feature data to an information processing terminal to display the virtual display object; A processing device for acquiring status information, including physiological sensor data, image data, and voice data, and for inferring the user's emotional and physical state based on the status information. A processing apparatus for generating prompt statements input to a generative artificial intelligence model using state summary information about the emotional state and the physical state, and generating response content of the virtual display object based on the response results obtained from the generative artificial intelligence model; A processing device for determining facial expression information and action information corresponding to the response content based on the response content, and sending the facial expression information and action information to the information processing terminal to control the display behavior of the virtual display object; A processing device for acquiring user operation information and behavior history information regarding the response content, and updating the generation conditions of the prompt statement and the response method of the virtual display object based on the operation information and the behavior history information.

[0761] (Note 2) The information processing system according to Appendix 1 is characterized in that, The processing device is configured to parse the prompt statement using a natural language processing algorithm, and generate control information for changing the visual attributes and behavioral characteristics of the virtual display object based on the user's preference information and emotional state.

[0762] (Note 3) The information processing system according to Appendix 1 is characterized in that, The processing device is configured to dynamically change the prompts input to the generative artificial intelligence model based on the user's emotional and physical state, and generate the response content of the virtual display object based on the response results obtained from the generative artificial intelligence model, thereby promoting the user's psychological state stability and behavior improvement through dialogue and notification processing with the user.

Claims

1. An information processing system, characterized in that, include: processor; The processor is configured to: parse prompts received from the user, generate feature data using prompts to indicate the appearance and personality of the avatar; send the generated feature data to the user's terminal to visualize the avatar on the user's terminal; and infer the user's emotional state, parsing prompts to generate avatar responses corresponding to the emotional state.

2. The information processing system according to claim 1, characterized in that, The processor is configured to use natural language processing technology to parse the prompt statement in order to understand the prompt indicating a change in avatar appearance and / or personality.

3. The information processing system according to claim 1, characterized in that, The processor is configured to generate corresponding avatar responses based on the user's emotional state and to alleviate the user's loneliness through interaction between the user and the avatar.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A