Information processing system

CN122802467APending Publication Date: 2026-09-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610247920.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-19
Filing Date
2026-03-02
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

然而,现有技术中普遍存在以下问题:其一,用户在创建虚拟角色时,通常需要分别操作图像生成系统、语音生成系统和聊天系统,各系统之间缺乏统一的控制和联动,导致用户体验分散、操作复杂;其二,现有聊天系统多以通用对话为主,缺少基于用户个体属性与情感状态的个性化虚拟角色形象和声音,无法为用户提供具有连续人格特征的虚拟陪伴;其三,尽管部分系统可以对文本进行简单情绪判断,但缺乏将文本与表情符号综合分析、并通过生成式AI自动生成能够给予用户安心感和情绪支持的信息的完整方案,因此难以有效缓解用户在压力、焦虑、孤独等情绪状态下的心理负担

Benefits of technology

[0021] "Presenting to the user" refers to the process of outputting system-generated images, text, or voice content to the user through a display screen, speaker, headphones, or other human-computer interaction interface, enabling the user to perceive and understand the content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802467A_ABST
    Figure CN122802467A_ABST
Patent Text Reader

Abstract

This invention provides an information processing system. The information processing system is characterized by comprising: a processor; wherein the processor is configured to: receive user-specified attribute information as input, generate prompt information to instruct an artificial intelligence model to generate a virtual character; generate prompt information to instruct an artificial intelligence model to generate speech based on the user's selection; parse text and emoticons received in the chat, analyze user emotions using natural language processing technology, and generate prompt information to instruct the artificial intelligence model to generate a message that provides reassurance based on the analysis results, and present the generated message to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology disclosed herein relates to an information processing system. Background Technology

[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot response to the user's speech.

[0003] With the development of generative artificial intelligence technology, users can create various forms of virtual content through image generation AI and voice generation AI. However, existing technologies generally suffer from the following problems: First, when creating a virtual character, users typically need to operate the image generation system, voice generation system, and chat system separately. The lack of unified control and linkage between these systems leads to a fragmented user experience and complex operation. Second, existing chat systems are mostly based on general dialogues, lacking personalized virtual character images and voices based on individual user attributes and emotional states, thus failing to provide users with virtual companionship featuring continuous personality traits. Third, although some systems can perform simple emotion judgments on text, they lack a complete solution that integrates text and emoticons for comprehensive analysis and automatically generates information that provides users with a sense of security and emotional support through generative AI. Therefore, it is difficult to effectively alleviate the psychological burden on users under emotional states such as stress, anxiety, and loneliness. Based on this, it is necessary to provide a system that can uniformly control virtual character generation, voice generation, and emotion analysis and comfort message generation to simplify the user operation process and enhance the immersiveness and emotional support effect of virtual companionship. Summary of the Invention

[0004] To address the aforementioned issues, this invention provides an information processing system comprising a processor configured to receive user-specified attribute information as input, generate prompts instructing an artificial intelligence (AI) model to generate a virtual character, thereby utilizing image-based AI to generate a virtual character image consistent with the user's personalized attributes. The processor is further configured to generate prompts instructing a voice-based AI model to generate speech based on the user's selection, thus configuring the virtual character with voice features matching the user's expectations. Further, the processor is configured to parse text and emoticons received in chat, analyze user emotions using natural language processing (NLP) technology, and generate prompts instructing the AI ​​model to generate reassuring messages based on the analysis results. This enables the system to automatically generate comforting and supportive responses tailored to the user's current emotional state and present the generated messages to the user. Through the above structure, this invention organically integrates virtual character image generation, virtual character voice generation, and sentiment analysis and comfort message generation based on chat content and emoticons into the same system. This enables the construction of personalized virtual characters driven by user attributes and intelligent companionship and emotional support oriented towards the user's emotional state. As a result, it effectively simplifies user operation steps, enhances the continuity and immersion of virtual interaction, and improves the user's psychological experience under negative emotions.

[0005] "System" refers to a collection of hardware and / or software devices or platforms used to perform functions such as virtual character generation, voice generation, sentiment analysis, and message generation.

[0006] A processor is a computing unit that can execute program instructions, process input data and output corresponding control signals or results, including but not limited to a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC) and other electronic circuits or combinations thereof with computing capabilities.

[0007] "User" refers to the entity that operates the system, which can be a natural person or other entity authorized to access the system through a terminal.

[0008] "Attribute information" refers to a set of parameters used to describe user preferences or to customize the characteristics of a virtual character for a user, including but not limited to information such as gender, age, hair color, eye color, body type, style preference, and voice preference.

[0009] "Prompt information" refers to instructional data or text used to drive or control the generative artificial intelligence model to perform a specific generative task, including but not limited to descriptions of generation conditions, constraints, and objectives given in natural language or structured parameters.

[0010] "Generative artificial intelligence models" refer to artificial intelligence models trained on large-scale data that can automatically generate content (including text, images, etc.) based on input prompts, including but not limited to large language models, image generation models, or combinations thereof.

[0011] "Virtual character" refers to a virtual human figure with specific appearance and / or behavioral characteristics generated by a generative artificial intelligence model based on user attribute information, which can be presented in the form of images, animations or other visualizations.

[0012] "Image generation AI" refers to an AI model or system that generates static or dynamic images based on input prompts, including but not limited to image generation engines based on Generative Adversarial Networks (GANs), diffusion models, or other generative models.

[0013] "Speech generation artificial intelligence" refers to artificial intelligence models or systems that generate speech signals based on input text and speech feature parameters, including but not limited to text-to-speech (TTS) models and deep learning-based speech synthesis models.

[0014] "Speech" refers to audio signals that can be perceived by the human ear and are generated by artificial intelligence. It usually exists in the form of digital audio data and can be played through speakers or headphones.

[0015] "Chat" refers to the information interaction process between users and systems based on natural language text. Users send messages by inputting text and / or emoticons, and the system returns corresponding responses in a dialogue format.

[0016] "Text" refers to the content composed of natural language characters that users input during the chat process, including sentences, phrases, words, etc.

[0017] "Emojis" are symbols used in chat to express or reinforce emotions, attitudes, or reactions, including but not limited to Emojis, icon-based emoticons, and other symbols that can express emotions.

[0018] Natural Language Processing (NLP) technology refers to a class of algorithms and models used to analyze, understand, and generate human natural language text, including but not limited to word segmentation, syntactic analysis, sentiment analysis, intent recognition, and text generation.

[0019] "User's emotions" refers to the psychological state or emotional tendency expressed or implied by the user at a specific point in time or in the context of a conversation, including but not limited to happiness, sadness, anger, anxiety, tension, relaxation, etc.

[0020] "Messages that provide reassurance" refer to text content generated by a generative artificial intelligence model based on the results of user sentiment analysis, used to soothe users' emotions, relieve stress, or provide emotional support. This includes, but is not limited to, comforting statements, encouraging statements, and empathetic responses.

[0021] "Presenting to the user" refers to the process of outputting system-generated images, text, or voice content to the user through a display screen, speaker, headphones, or other human-computer interaction interface, enabling the user to perceive and understand the content. Attached Figure Description

[0022] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.

[0023] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.

[0024] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.

[0025] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.

[0026] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.

[0027] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and head-mounted terminal according to the third embodiment.

[0028] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.

[0029] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.

[0030] Figure 9 This represents an emotion map that maps multiple emotions.

[0031] Figure 10 This represents an emotion map that maps multiple emotions.

[0032] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.

[0033] Figure 12This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.

[0034] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.

[0035] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation

[0036] Hereinafter, an example of an implementation of the system according to the present disclosure will be described with reference to the accompanying drawings.

[0037] First, let me explain the terminology used in the following instructions.

[0038] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.

[0039] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.

[0040] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.

[0041] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.

[0042] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.

[0043] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.

[0044] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.

[0045] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0046] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.

[0047] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.

[0048] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0049] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.

[0050] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.

[0051] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0052] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0053] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.

[0054] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.

[0055] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."

[0056] Existing technologies that utilize generative artificial intelligence models to generate virtual object images and related audio representations suffer from the following technical problems: (1) In traditional systems, the attribute information entered by users on the terminal is usually in a relatively scattered and low-level form, such as simple fields such as color, body shape, and clothing. The computer side usually constructs prompt statements by directly splicing text. There is a lack of a unified representation model that regularizes attribute information into higher-level conceptual feature information. As a result, the server lacks a structured semantic foundation when selecting generative artificial intelligence models, setting generation parameters, and reusing prompt statements in the future. It is difficult to achieve automated prompt statement optimization and high-quality visual representation generation.

[0057] (2) Existing image generation systems and dialogue systems are usually independent of each other: the image generation side only calls the image generation model based on a single prompt statement, while the dialogue side only performs natural language processing based on text. There is a lack of an integrated mechanism for unified format conversion, model type determination and model selection of prompt statements under the same computing architecture. It is impossible to dynamically switch between image generation model, text generation model and audio generation model according to task type and user status, resulting in low computing resource utilization efficiency and fragmented user experience.

[0058] (3) For text and symbol information entered by users through the chat interface, existing systems mostly only perform keyword matching or simple sentiment classification. They cannot directly convert the results of natural language processing and sentiment analysis into structured prompts that can be input into generative artificial intelligence models in a unified server process. It is difficult to automatically determine support strategies based on emotional state and generate personalized, reassuring response messages and corresponding audio performances, thus failing to effectively improve the quality of emotional support in human-computer interaction.

[0059] (4) In existing technologies, even if multi-round virtual object images and audio data can be generated, there is a lack of a historical management mechanism to establish a unified association record between feature information, prompts, visual representation data, audio data, and emotional states. The server has difficulty automatically generating corrective prompts, recommended candidate visual representations, or audio performances for the user based on existing records, and cannot continuously iterate and optimize the prompts in multi-round interactions, which limits the computer's adaptability and scalability in generative artificial intelligence applications.

[0060] (5) In terms of audio generation, traditional systems often require users to manually select parameters such as timbre and speech rate. They lack a mechanism to automatically deduce audio generation parameters based on the user's emotional state and the characteristics of virtual objects. The generated audio performance often lacks consistency with the visual representation and emotion analysis results, making it difficult to achieve semantic and emotional coordination and unity of multimodal output.

[0061] Therefore, a new system architecture and processing flow are needed. By introducing technologies such as attribute information regularization and structured generation of prompt statements on the server side, adaptive selection of generative artificial intelligence model types, generation of support strategies based on natural language processing and sentiment analysis, and historical association management of multimodal data and sentiment states, the performance and user experience of generative artificial intelligence applications can be improved from the perspective of computer system architecture and data processing flow.

[0062] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.

[0063] In this invention, the server includes: a device for receiving input information containing attribute information from a user terminal, regularizing the attribute information into feature information of a higher-level concept, thereby forming a data representation that can be uniformly processed within the server, and generating a prompt statement for instructing the generation of a visual representation based on the feature information; a device for converting the prompt statement into a form suitable for input into a generative artificial intelligence model, performing structured and formatted processing on the prompt statement, determining the type of the generative artificial intelligence model, and selecting between a generative artificial intelligence model for image generation and a generative artificial intelligence model for text generation based on the determination result, so as to achieve adaptive scheduling of models for different types of tasks; a device for inputting the converted prompt statement and image generation parameters into the selected generative artificial intelligence model for image generation to generate visual representation data reflecting the feature information, converting the visual representation data into image data that can be sent to the user terminal and sending it, thereby presenting a virtual object image corresponding to the feature information on the terminal side; and a device for determining whether there is a request for generating an audio representation based on the user's selection information, and if the determination result is positive, generating a prompt statement and audio generation parameters for input into a generative artificial intelligence model for audio generation, and transmitting the audio representation data to the generative artificial intelligence model for audio generation. A generative artificial intelligence model generates audio data and sends the audio data to a user terminal to achieve multimodal output in conjunction with visual representation. The device is used to obtain text and symbol information from user terminal text and symbol information received based on a session, analyze the text and symbol information using natural language processing and sentiment analysis techniques to classify the user's emotional state, determine a support strategy based on the emotional state, generate a prompt statement containing reassuring response content based on the support strategy, and input the prompt statement into a generative artificial intelligence model for text generation to generate a response message, thereby automatically... A device for constructing emotionally supportive dialogue content; a device for generating prompt statements for controlling a generative artificial intelligence model to generate audio data corresponding to the response message, generating the audio data when necessary, and prompting the response message and the audio data to a user terminal to achieve integrated text and audio feedback; and a historical management device for associating and recording the feature information, the prompt statements, the visual representation data, the audio data, and the emotional state, and for reusing prompt statement candidates or visual representation candidates to the user based on the records, thereby supporting iterative optimization and reuse of existing generated results.This allows for unified modeling and automated control of the expression of user attribute information, the selection and scheduling of generative artificial intelligence models, and the logic of multimodal data generation and emotional support at the computer system level. This improves the semantic consistency and emotional matching of the generated results, reduces the burden of manual adjustment for users, and enhances the efficiency and scalability of servers when handling multi-round interactions and historical data reuse, thereby substantially improving the computer technology performance in generative artificial intelligence applications.

[0064] "System" refers to an integrated information processing unit consisting of information processing devices such as servers, user terminals, and communication infrastructure, used to perform attribute information acquisition, prompt statement generation, generative artificial intelligence model invocation, and result output.

[0065] "User terminal" refers to an electronic device that allows users to input attribute information, view results, and perform interactive operations, including but not limited to mobile terminals, computer terminals, or other devices with display and input functions.

[0066] A "server" refers to a computing resource node equipped with a processor, memory, and network interface, used to perform program processing such as prompt statement generation, model selection, data processing, and history management.

[0067] "Attribute information" refers to the collection of various descriptive information specified by the user for the target virtual object or output content, including but not limited to low-level feature data such as color, shape, style, gender, expression, and background.

[0068] "Feature information" refers to the higher-level concept representation obtained by regularizing attribute information. It is used to describe the abstract characteristics of virtual objects or output content in a structured way so that the server can process them uniformly during the model selection and parameter setting stages.

[0069] "Prompt statements" refer to textual information used as input to generative artificial intelligence models, including descriptive statements about the visual representation of the target, text content, or audio content, as well as supplementary explanations related to the generation parameters.

[0070] "Generative artificial intelligence models" refer to computational models that are trained through machine learning and can automatically generate data such as images, text, or audio based on prompts. These include models for image generation, models for text generation, and models for audio generation.

[0071] "Generative AI Model for Image Generation" refers to a generative AI model that takes prompts and image generation parameters as input and outputs visual representation data, used to generate virtual object images or other image content.

[0072] "Generative AI model for text generation" refers to a generative AI model that takes prompts as input and outputs natural language text, used to generate response messages, explanatory text, or other language content.

[0073] "Generative AI model for audio generation" refers to a generative AI model that takes prompts and audio generation parameters as input and outputs audio data to generate speech, sound effects or other sound content.

[0074] "Visual representation data" refers to the internal image data representation corresponding to feature information output by a generative artificial intelligence model for image generation, including but not limited to pixel matrices, image files, or visual cache data.

[0075] "Image data" refers to image information that can be received and displayed by user terminals after being converted from visual representation data, including but not limited to bitmap data, compressed image files and their network transmission formats.

[0076] "Audio data" refers to sound information generated by generative artificial intelligence models for audio generation that can be played on user terminals, including but not limited to speech waveform data, compressed audio files and their network transmission formats.

[0077] “Text messages” refer to natural language text content submitted by users through chat or input interfaces, including sentences, phrases, words, and symbolic text.

[0078] "Symbolic information" refers to non-textual symbols that are related to a user's emotions or semantics, including but not limited to emoticons, punctuation marks, special symbols and their combinations.

[0079] Natural Language Processing (NLP) technology refers to a set of methods and algorithms used to process textual information, such as word segmentation, part-of-speech tagging, syntactic analysis, semantic understanding, and classification.

[0080] "Sentiment analysis technology" refers to the technology that uses natural language processing and statistical or learning models to calculate textual and symbolic information in order to determine the category and intensity of a user's emotional state.

[0081] "Emotional state" refers to the classification of a user's current emotions through sentiment analysis technology, including but not limited to positive, negative, neutral, anxious, excited, and other emotional categories.

[0082] "Support strategy" refers to a set of rules that guide the generation of response content and interaction methods that provide a sense of security or other psychological support, based on the coping principles determined by the emotional state.

[0083] "Response message" refers to natural language text output by a generative artificial intelligence model based on a support strategy, used to provide users with reassurance, advice, or instructions.

[0084] The "history management device" refers to a functional module implemented in the server that associates and records feature information, prompts, visual representation data, audio data, and emotional states, and provides users with prompts or visual representations based on these records.

[0085] "Image generation parameters" refers to the set of control parameters used when calling a generative artificial intelligence model for image generation, including but not limited to image resolution, number of generation steps, random seed, style intensity, and guidance coefficient.

[0086] "Audio generation parameters" refers to the set of control parameters used when calling the generative artificial intelligence model for audio generation, including but not limited to timbre type, speech rate type, pitch setting, and performance style type.

[0087] "Timbre type" refers to the category information used in audio generation to distinguish vocal characteristics, including but not limited to gender characteristics, age characteristics, and timbre style characteristics.

[0088] "Speed ​​type" refers to the parameter category used to set the playback speed of speech in audio generation, including but not limited to slow, normal, and fast, and their sub-levels.

[0089] "Expression style type" refers to the classification information used in audio generation to set the style of expressing emotions and tone, including but not limited to style categories such as gentle, lively, serious, and calm.

[0090] The embodiments of the present invention will be described based on the system architecture defined in the foregoing claims. In the following paragraphs, the server, terminal, and user are respectively the subjects performing different technical processes, so that the present invention can be reproduced and run in specific hardware and software environments.

[0091] In one implementation, the server is deployed as a centralized information processing node within a data center rack. The server includes a multi-core central processing unit (e.g., a general-purpose processor with a vector instruction set), a graphics processing unit (e.g., a parallel processing accelerator supporting CUDA), main memory, and high-speed solid-state storage media. The server runs application server software and deep learning inference frameworks (e.g., frameworks based on tensor operations libraries and automatic differentiation libraries) on top of an operating system (e.g., a Unix-like operating system), and pre-loads various generative artificial intelligence models, including generative artificial intelligence models for image generation, text generation, and audio generation.

[0092] In one embodiment, the terminal is a mobile computing device equipped with a display, a touch input component, and a wireless communication module, running a mobile operating system and having the graphical user interface application specific to this invention installed thereon. The terminal can also be a general-purpose computer terminal running browser software, in which case the browser executes front-end scripts to achieve interface presentation and network request sending.

[0093] In this implementation, users input attribute information through the terminal's graphical interface. Users select multiple fields on the terminal interface, such as color, hairstyle, eye shape, body type, clothing style, background environment, and art style. The terminal temporarily stores these fields in memory in a structured manner and performs simple validity checks locally, such as checking if a field is empty or if its value is within a preset enumeration range. Users can also directly input free descriptions in the natural language input box, such as: "Please generate a two-dimensional male character with long red hair, blue eyes, a muscular physique, wearing a black leather jacket, against a nighttime city street background, in a cyberpunk style." The terminal merges the structured fields with the free text to form candidate prompts, which are then displayed to the user for confirmation and modification.

[0094] After receiving attribute information from the terminal, the server constructs an intermediate data structure containing multiple key-value pairs in memory. The server then maps these keys and values ​​to its internal feature space. For example, the server maps various natural language expressions such as "red" and "red hair" to the standardized feature label "hair_color: red," and "muscular" and "robust" to "body_type: muscular." Using a predefined set of mapping tables and a hierarchical ontology structure, the server elevates the user-input attribute information to higher-level conceptual feature information, such as categories like "color features," "body type features," "style features," and "background features," and stores them in a unified format in the database. This feature regularization process enables the server to perform matching and optimization based on structured features when subsequently selecting generative AI models, setting generation parameters, and executing historical searches, rather than relying solely on simple string comparisons, thereby improving processing accuracy and efficiency.

[0095] After regularizing the feature information, the server uses natural language templates and rules to convert the feature information into prompts for inputting generative AI models. In one implementation, the server maintains different sets of prompt templates for different models. For example, for a generative AI model for image generation, the server uses templates that emphasize visual details and style tags, combining features such as "hair_color: red", "eye_color: blue", "style: anime", and "background: night_city_cyberpunk" to generate the text description: "cyberpunk style, night city street, anime male character, red long hair, blue eyes, muscularbody, wearing black leather jacket, highly detailed". This feature-based template generation method avoids semantic inconsistencies caused by the terminal simply concatenating strings, reduces noise in the prompts, and thus improves the stability of the subsequent generative AI model in visual feature mapping.

[0096] In another implementation, the server performs word segmentation and part-of-speech tagging on the user's freely input natural language description, and determines the feature category of key phrases using an internally trained classification model. The server inserts the identified phrases into the aforementioned feature information structure, and then merges them with the mapped higher-level concept features to form a more complete feature set. In this way, the server can utilize both structured form input and free text input simultaneously, improving its ability to parse complex descriptions.

[0097] After receiving the prompt, the server executes model type determination logic. The server maintains a decision module that selects a generative AI model for image generation, text generation, or audio generation based on the current task type flag, user request parameters, and keywords in the prompt. For example, if the server detects a task type of "visual generation" and the prompt includes style tags and resolution parameters, it selects an image generation model; if it detects a task type of "emotional response," it prioritizes a text generation model. This decision module is implemented using a simple decision tree and threshold rules, avoiding the uniform use of the same large model across all tasks and reducing unnecessary computational overhead.

[0098] In one specific embodiment, the server employs a generative artificial intelligence model for image generation based on an encoder-decoder architecture. The server first transforms the prompt into a fixed-length sequence of text feature vectors using a text encoder (e.g., a language coding network based on multi-head attention layers). Then, the server initializes a random noise tensor in the image latent space and progressively denoises it across multiple time steps. At each time step, the server calculates the residual based on the text features and the current noise state by calling a submodule composed of convolutional and attention networks, and updates the latent representation. Finally, the server maps the latent representation to a pixel-space image using a decoding network to obtain the visual representation data. This generation process is accelerated using matrix multiplication and convolution operations on a GPU, and the server balances quality and speed by controlling parameters such as the number of time steps, guidance coefficients, and resolution.

[0099] In another embodiment, the server invokes a generative AI model for text generation based on an autoregressive structure. The server encodes emotionally supportive prompts as sequences of words and maps them to a continuous vector space through embedding layers. The server iteratively computes the contextual representation of each position using a multi-layer self-attention network, then progressively predicts the probability distribution of the next word. The server employs a beam search or temperature sampling strategy to generate response message text with appropriate diversity and coherence. Because this model structure can simultaneously consider long-distance dependencies and local emotional cues, the server can generate statements that better meet the user's psychological needs based on emotional state, while ensuring semantic accuracy.

[0100] In the audio generation scenario, the server invokes a generative AI model for audio generation, which is composed of an acoustic model and a vocoder model. The server first converts the text to be read into a sequence of phonemes or subwords, and then calculates the corresponding acoustic feature sequence, such as a Mel spectrogram, using multi-layer convolutional or attention networks. Subsequently, the server invokes a vocoder based on an autoregressive convolutional or streaming generation structure to convert the Mel spectrogram into a time-domain speech waveform. Before invoking, the server automatically selects audio generation parameters such as timbre type, speech rate type, and performance style type based on the user's emotional state and the feature information of the virtual object. For example, it selects a combination of soft bass and a slightly slower speech rate when the user is experiencing negative emotions. The timbre and speech rate parameters are specifically reflected in different settings of the speaker embedding vector and time stretching factor within the network. Through this parameterized control, the server can generate diverse audio outputs that are consistent with visual representation and sentiment analysis results without retraining the model.

[0101] In this invention, the server incorporates a history management device to associate and record feature information, prompts, visual representation data, audio data, and emotional states. The server defines a record table in the database containing user identifiers, timestamps, feature vectors, prompt text, generation model identifiers, image file identifiers, audio file identifiers, and emotional tags. Each time the server completes a generation task, it writes a record. In subsequent requests, when the user wishes to adjust an existing virtual object, the server queries this record table and extracts the corresponding feature information and prompts. Based on the user's current modification intention, the server synthesizes historical prompts with new feature information to create a new prompt. For example, based on the original description of "long red hair, blue eyes, muscular physique," it automatically generates "While maintaining the overall style of the character, please change the hair color to silver and make the expression more aloof." This history-based prompt reconstruction helps maintain a stable character image across multiple interactions, while reducing repetitive user input and improving the utilization of computing resources.

[0102] The server utilizes a specially trained sentiment classification model for sentiment analysis. It merges the text and symbol information entered by the user in the chat interface into a unified text sequence, mapping emoticons and special symbols to predefined tags, such as "[SAD_FACE]" and "[EXCITED_MARK]". Subsequently, the server computes the representation of the entire text through an embedding layer and a multi-layer self-attention network, outputting probability estimates for various sentiment states via a classification head. The server combines rule analysis of the frequency of emoticons and punctuation marks; for example, it increases the weight of excitement when multiple exclamation marks appear consecutively, thus optimizing the classification results. Based on the final sentiment state, the server selects the corresponding support strategy template, then integrates the template with contextual information to form prompts for a generative AI model used for text generation, thereby generating reassuring, suggestive, or encouraging response messages. This processing chain avoids misjudgments caused by simple keyword matching, enabling the server to stably distinguish user emotional changes across different conversation rounds.

[0103] In this invention, the server divides the aforementioned modules into an input processing module, a feature regularization module, a prompt statement construction module, a model selection module, a generation and execution module, a sentiment analysis module, a history management module, and an output presentation module. Internally, the server passes intermediate data structures from one module to the next through message queues or function calls. Each module only depends on the structured data of the previous module and does not directly access the original user input text. This hierarchical structure allows the server to replace or upgrade downstream generative artificial intelligence models without modifying upstream modules, thereby enhancing the system's maintainability and scalability.

[0104] In terms of display output, the terminal decodes and presents the image and audio data returned by the server. The terminal uses an image rendering component to display virtual object images on the screen, and shows corresponding prompts or summary information next to the images, allowing users to intuitively understand the relationship between the generated results and the input features. When receiving audio data, the terminal decodes the audio stream through a multimedia playback engine and outputs it to the speaker. The terminal can also present emotional support response messages as text bubbles, and provide buttons for users to choose whether to play the corresponding comforting voice or generate a new character based on the response message.

[0105] When using the system, users can directly input natural language prompts, such as: "Please generate a cyberpunk-style anime male character with long red hair, blue eyes, and a muscular physique, wearing a black leather jacket, against a nighttime city street background." Users can also input prompts requesting audio and emotional support, such as: "I'm feeling a bit down today. Please generate an encouraging character and some comforting words, and add a gentle voice to these words." The server automatically selects an appropriate generative AI model and completes multimodal generation based on these prompts and attribute information.

[0106] This invention, by introducing specific processing steps such as feature information regularization, templated generation of prompt statements, adaptive selection of model types, and historical association management, enables the server to organize and process data in a structured manner internally. Compared to simple string concatenation and single model invocation, this significantly reduces the redundancy and uncertainty of prompt statements, and improves the semantic consistency and image detail matching accuracy of the generated results. Furthermore, by introducing automatic parameter selection based on emotional state and feature information in audio generation, this invention establishes an intrinsic correlation between audio performance, visual representation, and sentiment analysis results. This enables multimodal data generation under unified constraints within the computer, improving not only the user experience but also the overall utilization efficiency of computing resources.

[0107] The various embodiments of this invention can be modified. For example, the server can be replaced with a cluster of multiple distributed nodes. The task scheduling module can allocate image generation tasks to nodes with high-performance graphics processors, and text generation and sentiment analysis tasks to general-purpose computing nodes, thereby further improving processing throughput. Furthermore, the feature regularization module can use other types of ontology structures and word vector representations to support attribute information from more domains. As long as the core logic of attribute information regularization, structured prompt statements, model type selection, and historical association management is maintained on the server side, the technical effects of this invention can be achieved.

[0108] use Figure 11 The processing procedure is explained.

[0109] Step 1: The user enters attribute information on the terminal and generates an initial prompt statement.

[0110] In the terminal's graphical interface, users can input attribute information of the target object, such as hair color, eye color, body shape, clothing, background, style, and whether audio is required, through drop-down menus, radio buttons, sliders, and text input boxes.

[0111] The input to the terminal in this step consists of interface control values ​​and free text content generated by user operations.

[0112] The terminal assembles the values ​​of each control into a structured data record in memory, such as a set of key-value pairs, and concatenates the free text content with the structured attributes and fills in the template to generate an initial prompt statement in natural language.

[0113] The terminal outputs the following in this step: a set of structured attribute data and an initial prompt statement, which is then displayed on the interface for the user to confirm or modify.

[0114] Step 2: The terminal sends the attribute data and initial prompt statement to the server.

[0115] The terminal input in this step is the structured attribute data generated in step 1 and the initial prompt statement.

[0116] The terminal invokes the network communication module to encode the aforementioned data into a request message, which is then sent to the interface address specified by the server via the communication network. Before sending, the terminal performs serialization, necessary compression, and encryption processing on the data, and initiates the logic for waiting for a response after sending.

[0117] The terminal outputs the following in this step: a request data stream transmitted over the network to the server, containing user identifier, attribute data, and initial prompt statement.

[0118] Step 3: The server receives the request and builds an internal intermediate data structure.

[0119] The server's input at this step is a request message received from the network layer, which includes attribute data and an initial prompt statement.

[0120] The server, through the application server module, parses requests, extracts user identifiers, attribute fields, and initial prompts, and constructs a unified data object in memory, mapping all fields to an internal key-value structure. The server also records the request time and source information for log tracking.

[0121] The server output at this step is a standardized request data object containing original attribute information, original prompt statements, and metadata, which is used by subsequent processing modules.

[0122] Step 4: The server regularizes the attribute information into higher-level concept feature information.

[0123] The server's input in this step is the attribute information fields and initial prompt statement from the request data object obtained in step 3.

[0124] The server utilizes predefined mapping tables and ontology structures to perform lookup operations and rule matching on attribute values, merging diverse natural language expressions into unified feature labels. For example, the server maps "red hair" and "red hair" to "hair_color:red," and "muscular" and "robust" to "body_type:muscular." Simultaneously, the server combines the initial prompts with word segmentation and phrase recognition, supplementing any additional descriptions not yet entered through the form into new feature entries.

[0125] The server output in this step is a feature information structure containing multiple higher-level concept feature labels, which is used for subsequent prompt statement construction and model selection.

[0126] Step 5: The server generates structured prompts based on the feature information.

[0127] The server's input in this step is the feature information structure generated in step 4.

[0128] The server maintains template sets based on different model types, inserts feature labels into the corresponding templates, and generates more standardized prompt text through string replacement and sequential arrangement algorithms. For example, based on labels such as "hair_color:red", "eye_color:blue", "style:anime", and "background:night_city_cyberpunk", the server generates: "cyberpunk style, night city street, anime male character, red long hair, blue eyes, muscular body, wearing black jacket, highly detailed".

[0129] The server prioritizes features during the generation process (e.g., core features of the person are prioritized, followed by clothing, background, and style) to reduce ambiguity in the model's understanding.

[0130] The server output in this step is a structured prompt statement suitable for inputting a generative artificial intelligence model.

[0131] Step 6: The server determines the type of task to be generated and selects the type of generative artificial intelligence model.

[0132] The server's input in this step is the structured prompt statement from step 5 and the task flags in the request data (e.g., whether to generate an image, whether to generate a text response, whether to generate audio).

[0133] The server categorizes tasks based on decision rules: when the task flag indicates "visual generation" and the prompt contains image-related feature labels, the server labels the task as image generation; when the task flag indicates "emotional response," it labels it as text generation; and when the task flag indicates "generating speech," it labels it as audio generation. Based on the task type, the server selects the corresponding generative AI model for image generation, text generation, or audio generation from its internal model list.

[0134] The server output in this step is: a selected generative artificial intelligence model identifier and its corresponding model invocation configuration.

[0135] Step 7: The server generates visual representation data in the image generation task.

[0136] The server's input in this step is the structured prompt statement from step 5, the image generation parameters requested by the user (such as resolution, number of steps, and guidance coefficients), and the generative artificial intelligence model for image generation selected in step 6.

[0137] The server first invokes the text encoding sub-network to convert the prompt statement into a text feature tensor; then, it generates a random noise tensor in the latent image space as the initial state. In multiple time iterations, the server inputs the text features and the current noise state together into a multi-layer convolutional and attention network to calculate the noise residual and update the latent representation. This process is achieved through a combination of matrix multiplication, convolution, and non-linear activation operations. After iteration, the server invokes the decoding sub-network to map the latent representation into an RGB pixel matrix, forming the visual representation data.

[0138] The server output in this step is a high-dimensional image tensor or image file data corresponding to the feature information.

[0139] Step 8: The server converts the visual representation data into image data and sends it to the terminal.

[0140] The server's input for this step is the visual representation data obtained in step 7 and the network identifier of the user terminal.

[0141] The server uses an image encoding library to convert visual representation data into standard image file formats (such as PNG or JPEG) and generates thumbnail versions based on network bandwidth. The server writes the image file to storage media, obtains the file path or network access address, and then constructs a response message, encapsulating the image access path or image binary data in the response body.

[0142] The server's output in this step is: response data sent over the network to the terminal, which includes image data that can be displayed on the terminal or its access address.

[0143] Step 9: The terminal receives image data and displays virtual object images.

[0144] The terminal's input in this step is the response data returned by the server, including image data or image URLs.

[0145] The terminal receives the response via the network communication module. If it contains a URL, it further requests the image file; if it contains binary image data, it caches it locally. The terminal invokes the image rendering component to decode the image into a pixel buffer and draw it on the display screen. Simultaneously, the terminal can display brief text information next to the image, explaining the correspondence between the image and the user-input attributes.

[0146] The terminal's output in this step is: a virtual object image displayed on the screen, and the updated interface state.

[0147] Step 10: Users submit audio generation requests based on the generated results.

[0148] The user's input in this step is the virtual object image displayed on the terminal and the operation controls visible in the interface.

[0149] Users can click the "Generate Character Voice" button or enter a request in the text box, such as "Generate a self-introduction voice for this character, read in a gentle male voice," indicating their desire for the system to generate an audio performance corresponding to the current character. The terminal will then use this text as a new prompt, along with the current character's characteristic information and image identifier.

[0150] The terminal outputs the following in this step: a new round of request data containing audio generation requests, prompts, and character feature information.

[0151] Step 11: The server generates audio generation parameters and the text to be read aloud based on the audio request.

[0152] The server's input in this step is the audio generation request data sent by the terminal in step 10, including prompts and character feature information.

[0153] The server first uses a generative AI model for text generation to generate a suitable script for reading aloud, based on character traits and the user's desired tone. For example, it might say, "Hello everyone, I'm a bounty hunter guarding Night City. It's a pleasure to fight alongside you." The server then uses a sentiment analysis module to analyze the emotional cues in the user's current input, determining whether the user is in a depressed, excited, or neutral state. It then selects the appropriate voice type, speech rate, and performance style based on the character's traits. For instance, in a depressed mood, the server might choose a gentle, slightly slower male voice.

[0154] The server outputs the following in this step: a text to be read aloud and a set of audio generation parameters, which are used to drive the generative artificial intelligence model for audio generation.

[0155] Step 12: The server calls an audio generator to produce audio data using a generative artificial intelligence model.

[0156] The server's input in this step is the text to be read aloud generated in step 11, the audio generation parameters, and the selected generative artificial intelligence model for audio generation.

[0157] The server converts the read-aloud text into phoneme or word sequences and maps them to vector representations within the model via embedding layers. Combining audio generation parameters, the server injects corresponding timbre vectors, speech rate adjustment coefficients, and style control scalars into the acoustic model. The acoustic model then outputs acoustic features such as Mel spectra through multi-layer neural network computation. Subsequently, a vocoder network converts these features frame-by-frame into time-domain waveform data. Finally, the server encodes the waveforms into a compressed audio format file.

[0158] The server output at this step is a playable audio data file or audio data stream.

[0159] Step 13: The server sends the audio data and related text to the terminal.

[0160] The server's input for this step is the audio data generated in step 12 and the text content to be read aloud.

[0161] The server stores the audio file in a file system or object storage, generates an access link, constructs a response message including the audio access address and the text to be read aloud, and sends it to the terminal over the network. If network conditions permit, the server can also return the audio data directly as a binary stream.

[0162] The server output at this step is a response containing audio data and corresponding text content.

[0163] Step 14: The terminal plays audio and displays the dialogue in sync.

[0164] The terminal input in this step is the audio access address or audio data returned by the server and the corresponding text to be read aloud.

[0165] The terminal uses a multimedia playback component to load audio files or audio streams, and begins decoding and playback when the user clicks the play button. At the same time, it displays the text to be read as subtitles on the screen, and synchronizes the text and audio progress when necessary.

[0166] The terminal's output in this step is: voice output through the speaker and text subtitles displayed on the screen.

[0167] Step 15: Users enter text and symbols in the chat interface to express their emotions.

[0168] The user's input in this step is their current emotional state and the chat input box displayed on the terminal.

[0169] Users enter text and symbols in the chat input box, such as "I've tried many times today but I'm not satisfied, I'm a little annoyed :(", and click the send button. The terminal treats this combination of text and symbols as a new message entry.

[0170] The terminal outputs chat message data containing text and symbol information in this step and sends it to the server.

[0171] Step 16: The server performs sentiment analysis and generates sentiment support prompts.

[0172] The server's input in this step is the chat message data sent by the terminal in step 15.

[0173] The server uses a preprocessing module to convert emoticons and special symbols into predefined tags, transforming the entire message into a unified text sequence. Then, a sentiment classification model encodes and categorizes the text, obtaining sentiment labels and their probability distributions. Based on probability thresholds, the server determines the user's primary emotional state, such as "annoyed" or "depressed," and selects the corresponding support strategy from a strategy table. The server then embeds the user input, sentiment labels, and support strategies into a text generation template, constructing prompts for the generative AI model used in text generation. For example, "Please generate a short, comforting message in a soothing and encouraging tone, addressing the user's frustration after multiple unsuccessful attempts." The server output in this step is: an emotional support message and an emotional status label.

[0174] Step 17: The server generates an emotionally supportive response message and optionally generates corresponding audio.

[0175] The server's input in this step is the emotional support prompts and emotional status labels obtained in step 16.

[0176] The server invokes a generative AI model for text generation to generate a natural language response based on the prompts, such as, "I understand you're a little frustrated right now. Creativity sometimes does require multiple attempts. We can work together to adjust the prompts to make the result closer to your ideas. Please don't be discouraged." The server then decides whether to generate a voice version based on the emotional state: if voice support is needed, it re-executes the audio generation parameter selection logic, typically choosing a softer tone and a slower speaking speed, and invokes the generative AI model for audio generation to generate a corresponding comforting voice message.

[0177] The server outputs the following in this step: a text-based emotional support response message, and an optional emotional support audio message.

[0178] Step 18: The server returns emotional support results to the terminal and updates the history.

[0179] The server's input for this step is the emotional support response text generated in step 17, optional audio data, and the current user identifier and context information.

[0180] The server constructs a response message, packaging and sending it to the terminal along with the emotional support text, audio access address (if applicable), and emotional tags. Simultaneously, the server invokes the history management device to write the characteristic information, prompts, visual representation data identifiers, audio data identifiers, and emotional states from the current session into the database for future use in prompt candidate recommendation or visual representation candidate generation.

[0181] The server's output at this step is: an emotional support response sent to the terminal, and a historical record written to the database.

[0182] Step 19: The terminal displays emotional support content and allows users to regenerate it based on history.

[0183] The terminal's input in this step is the emotional support text, audio access address, and emotional tags returned by the server.

[0184] The terminal displays emotional support text in the chat interface and a "Play Comforting Voice" button based on whether an audio access address is available. The terminal can also retrieve historical prompts related to the current conversation from the server and present them as a clickable list. When the user clicks on a candidate, the terminal sends it as a new prompt to the server for generating new images or audio.

[0185] The terminal outputs the updated interface display status and the historical prompt statements that the user can directly access in this step, thereby supporting subsequent multiple rounds of generation operations.

[0186] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0187] In recent years, with the development of virtual spaces, online interactive environments, and human-computer dialogue technologies, users' demand for interacting with virtual representations through information terminals has been continuously increasing. However, existing virtual representation generation and dialogue support systems based on generative artificial intelligence models still suffer from the following computer technology-related limitations, making it difficult to improve the overall system performance and user experience.

[0188] First, existing systems generally adopt a simple one-time mapping method from attributes to images. The server usually just concatenates the appearance attributes input by the user into a fixed-format text description and inputs it into the image generation model. This method does not perform standardized modeling of attribute information with higher-level concepts, resulting in: (1) redundant and inconsistent attribute expressions, which are not conducive to reuse between different tasks; (2) difficulty in sharing attribute semantics among multiple generative artificial intelligence models, thereby increasing the complexity and resource consumption of server-side model calls; (3) difficulty in performing multi-round, continuous, and fine-grained updates to the virtual representation in the same session.

[0189] Second, in existing technologies, visual generation and speech generation are mostly independent modules, and the server lacks a unified mechanism for generating prompts and managing control parameters. The server often constructs inconsistent input descriptions for the image model and the speech model, making it impossible to establish a consistent system-level representation between visual features (such as physique, appearance color, expression, and clothing) and speech features (such as voice type, sense of age, speech rate, and emotional expression). This inconsistency will result in: (1) a disconnect between the appearance and voice style of the virtual representation, making it difficult to form a unified personality; (2) the server needs to maintain multiple sets of incompatible description templates and business logic, increasing maintenance costs and the probability of errors; and (3) consuming more computing and storage resources when generating multiple models collaboratively.

[0190] Third, existing systems often rely on simple keyword matching or sentiment labeling for emotion perception and dialogue support, and frequently decouple sentiment analysis from generative artificial intelligence models. The server cannot structurally embed the user's emotional state into prompts and model control parameters. The results are: (1) The system's response to the user's real emotions is slow or one-sided, and it cannot dynamically adjust the behavior and speech of the virtual representation; (2) Reassurance or psychological support responses are usually based on fixed templates, lacking specificity and continuity, making it difficult to improve the quality of human-computer interaction; (3) The server cannot fully utilize natural language processing results to perform fine-grained control over the downstream generation process, thus limiting the effectiveness of generative artificial intelligence models.

[0191] Fourth, existing virtual representation generation systems have technical deficiencies in terms of "real-time performance" and "continuous updates". Many systems only support users to set virtual images once, and the server does not dynamically update the results after they are generated. Even if updates are supported, users often need to re-enter the session or reload the scene, resulting in: (1) fragmented server-side state management and a lack of a continuous generation mechanism for the same session; (2) inability to achieve high-frequency iterative image and voice generation based on changes in user input; and (3) the inability to uniformly manage the session context, resulting in resource waste and response delays.

[0192] In summary, in computer implementations related to virtual representations, there is still a lack of a technical solution that can uniformly manage attribute information, emotional information, and prompts on the server side, and can efficiently and collaboratively generate visual information, speech information, and emotional support messages among multiple generative artificial intelligence models using a single higher-level concept representation, while supporting real-time and continuous updates based on user input.

[0193] Therefore, the objective of this invention is to provide a system and its computer implementation that can perform higher-level conceptual modeling of user attribute information, voice attribute information, and emotional information on a server, and collaboratively call multiple generative artificial intelligence models based on a unified prompt statement generation and control parameter management mechanism, so as to improve the consistency, real-time performance, and resource utilization efficiency of virtual representations in terms of visual, voice, and dialogue support, thereby improving the performance and scalability of human-computer interaction systems from a computer technology perspective.

[0194] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.

[0195] In this invention, the server includes: a device for acquiring attribute information input by a user through an information terminal and performing higher-level concept normalization processing; generating prompt statements to instruct a generative artificial intelligence model to generate visual information of a virtual representation; a device for generating voice generation control parameters or voice prompt statements based on user-selected voice attribute information and the attribute information; generating prompt statements to instruct the generative artificial intelligence model to generate voice information; a device for performing natural language processing parsing on text information and emoticon information received from the information terminal in a chat format and extracting emotional information; a device for constructing reassuring message candidates based on the emotional information and the attribute information and generating prompt statements to instruct the generative artificial intelligence model to generate the message candidates as natural language messages; a device for providing the virtual representation visual information, the natural language messages, and the voice information to the information terminal in real time for display and playback in a virtual space; and a device for sequentially acquiring changes made by the user to the attribute information and / or the voice attribute information and continuously updating the virtual representation visual information and voice information in time by re-executing the prompt statement generation processing and the generative artificial intelligence model generation processing. This allows for the coordinated control of multiple generative artificial intelligence models on the server side through a unified mechanism for representing higher-level concepts and generating prompts. This enables the integrated and real-time generation of virtual representations' appearance, voice, and emotional support messages, improving the efficiency of computing resource utilization and enhancing the system's interactive consistency and response performance in virtual space. Consequently, it addresses the technical problems of scattered attribute expressions, difficulty in integrating emotional information, and difficulty in continuously updating virtual representations in existing systems from a computer technology perspective.

[0196] A “system” refers to a computer implementation consisting of at least one server and at least one information terminal connected through a communication network, used to perform virtual representation generation, speech generation, and dialogue support processing.

[0197] A "server" is a computing device equipped with a processor and memory and running program instructions to perform core computing tasks such as attribute information processing, prompt statement generation, model calling, and data distribution.

[0198] "Information terminal" refers to an electronic device operated by the user and used to send and receive data with the server, to obtain user input, display virtual representations, and play voice information.

[0199] "Attribute information" refers to a set of various parameters used to describe the appearance or characteristics of a virtual representation, including but not limited to body information, appearance color information, expression information, and clothing information.

[0200] "Speech attribute information" refers to a set of various parameters used to describe the speech characteristics of a virtual representation, including but not limited to voice type information, age perception information, speech rate information, and emotional expression information.

[0201] "Higher-level concept normalization processing" refers to the process of abstracting and normalizing the specific attributes of user input, and converting them into a high-level semantic representation suitable as a general input for generative artificial intelligence models.

[0202] "Generative artificial intelligence models" refer to artificial intelligence models that are trained based on machine learning and can automatically generate output data, including images, text, or speech, based on input prompts or control parameters.

[0203] "Generative AI model for image generation" refers to a generative AI model aimed at generating image data, which can generate image data representing the appearance of a virtual representation based on prompts.

[0204] "Generative AI model for speech generation" refers to a generative AI model aimed at generating speech data, which can generate speech signals based on text and speech attribute information.

[0205] "Prompt statements" refer to text descriptions or instructional statements generated by the server based on attribute information, voice attribute information, or emotional information, which are used as input to generative artificial intelligence models.

[0206] "Virtual representation" refers to a digital object used to represent a user or system entity in virtual space, including characters presented as images or three-dimensional images.

[0207] "Virtual representation visual information" refers to image data or data related to image display generated by generative artificial intelligence models for displaying the appearance of virtual representations on information terminals.

[0208] "Physical information" refers to a set of attributes used to represent the body shape or body type characteristics of a virtual representation, including but not limited to abstract descriptions such as body proportions and body size.

[0209] "Appearance color information" refers to the set of attributes used to represent the appearance color characteristics of a virtual representation, including but not limited to abstract descriptions such as hair color, eye color, and main color tone of clothing.

[0210] "Face information" refers to a set of attributes used to represent the facial expressions or emotional states of a virtual representation.

[0211] "Clothing information" refers to a set of attributes used to represent the type, style, or matching characteristics of clothing worn by a virtual representation.

[0212] "Sound quality attribute information" refers to a set of abstract parameters used to comprehensively represent the auditory characteristics of speech, including voice type information, age perception information, speech rate information, and emotional expression information.

[0213] "Speech type information" refers to attributes used to abstractly represent the characteristics of speech type, including but not limited to gender and timbre category.

[0214] "Age perception information" refers to an abstract attribute used to represent the age-related characteristics of speech in terms of auditory perception.

[0215] "Speed ​​information" refers to an attribute used to indicate the speed of speech playback or the pace of speaking.

[0216] "Emotional expression information" refers to a set of abstract attributes used to represent the type and intensity of emotions conveyed in speech.

[0217] “Voice information” refers to digital voice data generated by a generative artificial intelligence model for playback on information terminals.

[0218] "Text message" refers to the sequence of natural language characters that a user enters and sends through a chat interface on a messaging terminal.

[0219] "Emoji information" refers to graphic symbols or emoticon icons that users send during chat to express emotions or intentions.

[0220] "Emotional information" refers to structured data obtained by the server based on text and emoji information, used to categorize and represent the user's emotional state and intensity.

[0221] "Message candidates" refers to a collection of multiple text contents or their abstract expressions generated by the server based on emotional and attribute information, used to provide users with a sense of security or emotional support.

[0222] "Natural language messages" refer to text content expressed in natural language, generated by generative artificial intelligence models based on message candidates, and presented to users.

[0223] "Real-time delivery" refers to the process by which the server outputs data to the information terminal within a predetermined time constraint after receiving input or generating results, so that the user can perceive a continuous, low-latency response during the interaction.

[0224] "Continuously updated state in time" means that as user input or attribute changes, the server can call the generative artificial intelligence model multiple times and output new results, so that the visual and voice information of the virtual representation is continuously and periodically refreshed during the same session.

[0225] In this embodiment, the server operates as the core computing node, including a multi-core central processing unit, a graphics processing unit, a large-capacity random access memory, and non-volatile storage devices. The server runs multiple software modules on an operating system, including server programs for network communication, a runtime environment for deep learning inference, a database system for data management, and a system management program for logging and monitoring. For deep learning inference, the server uses a deep learning framework and a model inference engine, and executes generative artificial intelligence models on the graphics processing unit.

[0226] In this embodiment, the terminal operates as a user interaction device, which may be a smartphone, tablet computing device, or head-mounted display device. The terminal includes a touch screen, an audio playback device, a microphone, a graphics processing unit, and a network communication module. The terminal runs a client application or a web application on an operating system. The client application includes a user interface module, a rendering module, a network communication module, and a local caching module. The terminal establishes an encrypted communication channel with the server through the network communication module.

[0227] In this embodiment, the user interacts with the server through a terminal. The user inputs and selects attribute information and voice attribute information related to the virtual representation on the terminal interface, and converses with the virtual representation using text and emoticons. The user can see the visual information of the virtual representation generated by the server on the terminal interface and hear the voice information played by the terminal.

[0228] After receiving the attribute information input by the user, the server performs superordinate concept normalization on the attribute information. The server pre-stores an attribute vocabulary and mapping rules in its storage device. The attribute vocabulary maps a large number of specific descriptions to a finite set of superordinate concept labels; for example, it maps "blue hair," "dark blue hair," and "azure hair" to the superordinate concept "hair color: blue." During processing, the server uses hash tables or prefix trees to match strings, converting the user input into a standardized set of labels. This normalization process reduces the fragmentation of the attribute space, allowing for the use of a unified feature representation when subsequently calling different generative AI models.

[0229] After normalizing the attribute information, the server generates prompts for image generation. The server pre-stores prompt templates in its storage device. These templates use natural language to abstractly represent the virtual representation's body structure, appearance, color, expression, and clothing information. The server uses a string template replacement algorithm to insert normalized attribute tags into specified placeholders to generate complete prompts. For example, when a user specifies "hair color: blue, eye color: green, body type: slim, style: futuristic, personality: cheerful" in the terminal, the server generates the following prompt: “Generate an avatar with blue hair, green eyes, and a slim body type in a futuristic sci-fi style, with a cheerful personality.” Based on the prompt statement, the server can further add technical control descriptions, such as desired lighting, composition, and resolution, to improve the quality of the generated image. Internally, the server stores the prompt statement and its corresponding attribute tags in structured data format for subsequent log analysis and model optimization.

[0230] The server invokes a generative artificial intelligence model for image generation. In this embodiment, the server can use a diffusion-type generative model, which consists of a text encoding subnetwork, a noise prediction subnetwork, and an image decoding subnetwork. When using this model, the server first uses the text encoding subnetwork to encode the prompt statement into a high-dimensional semantic vector. The text encoding subnetwork can employ a transformer structure. During the backward inference phase of the diffusion process, the server executes the noise prediction subnetwork multiple times to progressively denoise the latent space from high-noise states. The noise prediction subnetwork can employ a convolutional neural network with skip connections or a network structure based on a multi-head attention mechanism. After the denoising iteration is completed, the server uses the decoding subnetwork to restore the latent representation to an image pixel matrix. The server performs the above tensor operations in parallel on the graphics processing unit to improve inference speed.

[0231] After generating the image, the server compresses the high-precision floating-point pixel matrix into an image file using an image encoding library. Internally, the server uses a caching system to record the image file's path or identifier and stores image metadata associated with the session identifier in a database. The server then sends the image data to the terminal via a network interface. Upon receiving the image data, the terminal uses its graphics rendering module to load the image as a texture and displays the virtual representation in a two-dimensional interface or three-dimensional scene, thus presenting the visual information of the virtual representation on the display device.

[0232] In speech generation, the server generates voice prompts or control parameters based on user-selected voice attribute information. The server pre-stores a voice attribute mapping table in its storage device, mapping specific descriptions such as "female voice," "young," "gentle," and "slow" to acoustic control parameters, including reference pitch, reference speech rate, pitch range, and energy curve. The server constructs input for the generative AI model used for speech generation; the input can include text content and control parameters. When using a sequence-to-sequence-based speech generation model, the server employs a text encoder and acoustic decoder structure, mapping text features to time-series acoustic features through an attention mechanism, and then converting the acoustic features into waveform signals through a vocoder network. When training this type of model, the server uses speech reconstruction error as a loss function, updating network weights through backpropagation and gradient descent to minimize the reconstruction error on the training set.

[0233] When generating speech based on personality and context, the server can construct the following speech prompts for the speech generation model: “Generate a gentle and cheerful young female voice for the following text.” The server inputs the prompt statement along with the text to be read into the speech generation model, causing the model to adjust its acoustic features according to the control description when generating speech. After receiving the speech data returned by the server, the terminal buffers and decodes the audio stream through the operating system's audio interface and plays the speech on the audio output device, presenting it in conjunction with the image display of the virtual representation.

[0234] The server employs a natural language processing model and an emoji parsing algorithm for sentiment analysis. It pre-stores a set of sentiment tags and a trained text sentiment classification model in its storage device. The server converts user-sent text into a word segmentation sequence, which is then transformed into a feature vector sequence through an embedding layer and a transformer encoder. Based on the encoded sequence, the server generates an overall sentence vector through pooling operations and outputs the sentiment category (e.g., happy, sad, tense, tired) and corresponding confidence score using a classification layer. Simultaneously, the server parses emoji information, converting them into sentiment weight vectors according to a predefined mapping table. The server then performs a weighted synthesis of the text sentiment results and emoji weights to obtain structured sentiment information. This structured sentiment information includes sentiment type, intensity, and confidence score, and serves as a control signal for downstream generation processing.

[0235] When generating reassuring messages based on emotional information, the server utilizes a combination of templates and a generative artificial intelligence model. First, the server selects an appropriate language style and structural template based on the emotional type; for example, it uses a structure combining encouragement and empathy in the context of "mild fatigue." The server then constructs prompt statements containing emotional descriptions and user attributes for the generative AI model and inputs them into the text generation model to generate a natural language message. The server can use the following forms when constructing prompt statements: “Generate a supportive and comforting message for a user who feels abit tired but still positive, considering the user has a cheerful personality.” The server uses a unified prompt generation module to express similar overarching concepts for three types of generative AI models: image generation, speech generation, and text generation. This reduces the complexity of manual adaptation between multiple models. Internally, the server shares standardized attribute and sentiment information, ensuring semantic consistency when different models use the same user session data.

[0236] The server employs a modular data structure throughout the system. It uses structured objects or records in memory to represent user sessions. This session structure includes a set of attribute tags, a set of voice attributes, emotional information, an identifier of the most recently generated result, and a history of prompts. The server persists some session data in the database using session identifiers to support session recovery and long-term interactions. Each time the server receives updated attributes or dialogue content from the user, it updates the fields in the corresponding session structure and minimizes the scope of recomputation based on the changed parts. It only re-infers the affected generative AI models, avoiding full recalculation of all models, thereby reducing computational load and response time.

[0237] The terminal performs input acquisition, result display, and simple preprocessing tasks on the user side. When the user inputs text and selects an emoji, the terminal encapsulates the data into a structured message and sends it to the server via network protocols. When receiving visual information, natural language messages, and voice information from the server, the terminal caches them locally to ensure smooth playback and display even under network fluctuations. When the user adjusts attributes or voice parameters, the terminal immediately sends an update request to the server.

[0238] When using this system, users can modify the appearance and voice attributes of the virtual representation multiple times consecutively. Each time the server receives a modification, it regenerates the prompts based on the normalized attribute information and invokes the relevant generative artificial intelligence model, thereby continuously updating the image and voice of the virtual representation. Users perceive these changes in real time through their terminals, without needing to restart the session or reload the virtual space scene.

[0239] The technical solution in this embodiment not only achieves personalized generation of virtual representations but also brings multiple improvements at the computer technology level. The server reduces redundancy in the attribute expression space and improves the reusability of model inputs by standardizing higher-level concepts and using a unified prompt generation mechanism. This allows the same set of attribute representations to be reused across multiple generative AI models, reducing memory usage and parameter conversion overhead. By structuring emotional information and inputting it as control signals to the text and speech generation modules, the server achieves multimodal collaborative generation based on unified emotional representation, improving response accuracy and consistency. In terms of data structure design and incremental update strategies, the server reduces the number of repeated inferences through session structures and partial recalculation mechanisms, thereby shortening generation latency and reducing the load on the graphics processing unit.

[0240] The server employs specific learning methods and optimization strategies during model training and inference. When training the image generation model, it uses a combination of reconstruction loss and contrast loss, leveraging large-scale image-text pairs for supervised learning to improve the accuracy of text-to-image mapping. When training the speech generation model, it uses mean squared error and log-likelihood loss to ensure the generated acoustic features closely approximate real speech. When training the sentiment analysis model, it uses cross-entropy loss and supervises multi-category sentiment labeling. During training, the server performs data augmentation on the input data, such as synonym substitution for text, random image cropping, and duration stretching or noise addition for speech, to improve model robustness. During inference, the server utilizes batch processing and multi-stream parallelism strategies for graphics processing units to merge and execute multiple user requests, thereby improving overall throughput.

[0241] In this implementation, the server technically addresses issues in traditional systems such as scattered attribute representations, inconsistencies between visual and speech styles, and the difficulty in integrating emotional information into the generation process. Because the server uses unified higher-level conceptual modeling for attribute information and shares this modeling result across multimodal generation tasks through prompts, it reduces repetitive parsing and adaptation logic in the computational path, thus improving computational efficiency and response speed. Furthermore, by using structured emotional information as the generation control signal, rather than simple template replacement, the server can provide more refined and continuous emotional support with the same hardware resources, thereby technically improving the relevance of the generated results and the consistency of the user experience.

[0242] The system of this invention can be applied to various virtual spaces and interactive scenarios. When a terminal uses this system in a virtual reality device, it can map the visual information of the virtual representation generated by the server onto the three-dimensional environment and drive the lip-sync animation of the three-dimensional character based on the voice information returned by the server. In this scenario, the server utilizes real-time inference and incremental update mechanisms to enable the virtual reality system to maintain a high frame rate and audio-visual synchronization even in bandwidth-limited and latency-sensitive network environments, thus demonstrating the technical effectiveness of this invention in actual device control and resource scheduling. Through the above-described configuration, the server transforms this system from a simple automation of manual operations into a system specifically designed for the input structure, attribute management methods, and multi-model collaborative strategies of generative artificial intelligence models, thereby achieving coordinated optimization of data representation, algorithm flow, and resource utilization within the computer.

[0243] use Figure 12 The processing procedure is explained.

[0244] Step 1: Users launch the application on the terminal and select the virtual representation editing function.

[0245] Users can select or input appearance-related attribute information and voice-related voice attribute information through clicking, swiping, and input operations on the terminal interface.

[0246] The terminal uses the current session identifier as the key and the attribute information and voice attribute information as the value, and temporarily stores them in the terminal's local memory.

[0247] The terminal packages the above information into a structured request message as the output of this step and prepares to send it to the server over the network.

[0248] The input for this step is the user's interactive operation and the existing session identifier. The output for this step is a structured request message containing attribute information and voice attribute information.

[0249] Step 2: The terminal sends structured request messages to the server through the communication module.

[0250] Before sending, the terminal serializes and compresses the message to reduce network transmission load, and appends a session identifier and timestamp to the message header.

[0251] After receiving the message at the communication interface, the server writes the message to the receive buffer and decodes the raw binary data into an internal structure.

[0252] The input for this step is a structured request message generated by the terminal, and the output for this step is a parsable attribute information and voice attribute information data structure on the server side.

[0253] Step 3: The server performs higher-level concept normalization processing on the received attribute information.

[0254] The server uses a pre-stored attribute vocabulary and mapping rules to map user-provided specific descriptions (such as "long blue hair" and "bright green eyes") to standardized labels (such as "hair color: blue", "eye color: green", and "body type: slim").

[0255] The server performs normalization operations on each field through string matching, hash lookup, and rule judgment, transforming the original text into an enumerated value or tag vector with internal uniform encoding.

[0256] The server outputs the normalized attribute tag set and stores it in the session data structure associated with the current session identifier.

[0257] The input for this step is the raw attribute information parsed by the server, and the output for this step is a standardized set of attribute tags.

[0258] Step 4: The server generates prompts for image generation based on the normalized attribute tags.

[0259] The server reads the natural language template from the storage device and uses a string replacement algorithm to replace placeholders such as "physical information", "appearance color information", "expression information", and "clothing information" with the corresponding tag text.

[0260] The server combines these parts to generate a complete natural language sentence, for example: “Generate an avatar with blue hair, green eyes, and a slim body type in a futuristic sci-fi style, with a cheerful personality.” The server stores the prompt statement as text data in the session structure and uses it as input to the image generation module.

[0261] The input for this step is a standardized set of attribute labels, and the output is the text of the prompt statement for image generation.

[0262] Step 5: The server inputs prompts for image generation into the generative artificial intelligence model for image generation.

[0263] The server uses a text encoding subnetwork to convert the prompt statement into a high-dimensional text feature vector, and then uses this vector as a condition in the diffusion generation subnetwork to perform multiple rounds of denoising iteration on the potential noise vector.

[0264] In each diffusion step, the server predicts the noise component based on the current latent representation and text features, and updates the latent representation until convergence yields a low-noise latent image vector.

[0265] The server uses a decoding subnetwork to map potential image vectors into pixel matrices and encodes the floating-point matrices into an image file format.

[0266] The input for this step is the prompt text for image generation, and the output for this step is image data representing the visual information of the virtual representation.

[0267] Step 6: The server sends the generated image data to the terminal.

[0268] The server includes image data or image resource addresses in the network response, and appends a session identifier and image version number to the response header.

[0269] After receiving the response, the terminal decodes the image data into a texture or bitmap object and renders and displays it in the corresponding area of ​​the user interface, allowing users to intuitively view the appearance of the virtual representation.

[0270] The input for this step is the image data already generated on the server side, and the output for this step is the virtual representation image displayed on the terminal screen.

[0271] Step 7: The server generates voice control parameters or voice prompts based on voice attribute information.

[0272] The server uses a speech attribute mapping table to convert labels such as "female," "young," "gentle," and "slow" into numerical control parameters, such as reference pitch, target speech rate, energy range, and emotion weight vector.

[0273] When text-based control is required, the server generates voice prompts such as "Generate a gentle and cheerful youngfemale voice for the following text." and saves them along with the control parameters.

[0274] The input for this step is standardized voice attribute information, and the output for this step is voice generation control parameters and / or voice prompts.

[0275] Step 8: Users can chat with virtual representations on the terminal.

[0276] Users enter text in the chat input box and can select or add emojis, such as " " "wait.

[0277] The terminal organizes the text string and the list of emojis into a structured dialogue message, attaches the current session identifier, and sends it to the server.

[0278] The input for this step is the text and emoji actions performed by the user on the terminal, and the output for this step is the dialogue message data sent to the server.

[0279] Step 9: The server performs sentiment analysis on the conversation messages.

[0280] The server segments the text information into words or sub-words and encodes it into a sequence of feature vectors using an embedding layer and a transform encoder; at the same time, the server converts emojis into emotion vectors according to a predefined mapping table.

[0281] The server aggregates and classifies text features through a sentiment classification subnetwork to obtain initial sentiment categories and confidence levels. Then, it performs weighted synthesis with facial expression sentiment vectors to obtain the final structured sentiment information.

[0282] The server stores the emotion category (such as "mildly tired", "happy", "sad") and intensity (numerical score) as emotion information in the session structure.

[0283] The input for this step is dialogue message data containing text and emojis, and the output for this step is structured sentiment information.

[0284] Step 10: The server generates reassuring message candidates based on emotional information and attribute tags.

[0285] The server selects the corresponding language style and tone template based on the emotion type. For example, it selects a sentence structure that includes empathy and encouragement for "mild fatigue".

[0286] The server writes the user's personality tags and current virtual representation settings into the template, and constructs prompts for text generation, such as: “Generate a supportive and comforting message for a user who feels abit tired but still positive, considering the user has a cheerful personality.” The server uses a generative artificial intelligence model to generate one or more natural language message candidates from the input text of the prompt statement.

[0287] The input for this step is a set of structured sentiment information and attribute labels, and the output for this step is one or more natural language message candidates expressing reassurance.

[0288] Step 11: The server selects or post-processes reassuring natural language messages and prepares speech to generate text.

[0289] The server scores multiple message candidates based on length, sentiment intensity matching degree, and keyword coverage, and selects the message with the highest score as the final reply text.

[0290] The server performs simple post-processing on the selected text, such as removing duplicate phrases and correcting punctuation, and uses the results as input text for speech generation.

[0291] The input for this step is multiple natural language message candidates, and the output for this step is the final response text used for speech generation and terminal display.

[0292] Step 12: The server calls a speech generation function to use a generative artificial intelligence model to generate speech information.

[0293] The server inputs the final response text along with the voice control parameters or voice prompts generated in step 7 into the voice generation model. The text encoder obtains the text feature sequence, and the acoustic decoder generates the corresponding acoustic feature sequence.

[0294] The server then uses a vocoder network to convert the acoustic features into waveform data, generating a digital audio stream.

[0295] The server encodes and compresses the audio stream to generate audio format files or data packets suitable for network transmission and terminal playback.

[0296] The input for this step is the final response text, voice control parameters, and / or voice prompts. The output for this step is playable voice data.

[0297] Step 13: The server will then send the final response text and voice data to the terminal.

[0298] The server carries natural language message text and corresponding audio data or audio resource address in the response. After receiving the response, the terminal hands it over to the text display module and the audio playback module for processing, respectively.

[0299] The terminal displays the reply text in the chat interface, and at the same time calls the audio interface to decode and play the voice, so that the user can see the reply content visually and perceive the sound performance consistent with the virtual representation.

[0300] The input for this step is the reply text and voice data generated on the server side, and the output for this step is the text display and voice playback effect on the terminal side.

[0301] Step 14: Users decide whether to continue editing attributes or continue the conversation based on the current display and playback results.

[0302] If a user is not satisfied with the appearance or sound of a virtual representation, the user can modify the attribute information or voice attribute information in the terminal interface.

[0303] After detecting an attribute change or new dialogue input, the terminal generates a new request message and sends it to the server. The server then repeats the aforementioned steps to incrementally update the images, text, and voice.

[0304] The input for this step is the user's subjective feeling about the current presentation effect and new interactive operations. The output of this step is a new round of data processing and generation triggered by new request messages.

[0305] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.

[0306] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."

[0307] With the development of virtual object generation technology and speech generation technology based on generative artificial intelligence models, existing systems still have many shortcomings in efficiently and accurately mapping user intentions into control information that can be understood by generative artificial intelligence models.

[0308] First, existing solutions typically require users to directly set a large number of low-level speech parameters (such as specific numerical values ​​for pitch, speech rate, timbre, etc.), or to describe them using general textual instructions that are difficult for the model to understand. This makes the parameter setting process complex, difficult for users to operate intuitively, and results in a large deviation between the generated results and the user's subjective intentions, thereby reducing the human-computer interaction experience.

[0309] Second, existing systems in chat scenarios mostly perform only surface analysis of text content and lack a mechanism for joint emotion analysis of textual and symbolic information (such as emoticons, punctuation patterns, etc.). They cannot start from high-level emotional states and automatically generate response content with a soothing or empathetic effect, thus showing significant shortcomings in improving users' emotional experience and psychological security.

[0310] Third, in existing technologies, instruction generation, model invocation, and result feedback are often separate processing flows. There is a lack of a systematic framework that unifies and abstracts multi-source information such as user attribute information, voice parameters, and emotional state into higher-level feature information, and automatically drives generative artificial intelligence models through prompt statements. There is also a lack of a closed-loop optimization mechanism that iteratively updates control information and adaptively corrects prompt statements based on user evaluation information of the generated speech reproduction results. As a result, the system cannot achieve self-learning adjustment and performance improvement of the generation process within the computer.

[0311] Fourth, existing implementations typically only encapsulate general model interfaces in terms of system architecture. They do not fully utilize the optimization of prompt generation, parameter abstraction, model selection, and invocation processes at the computer system level to improve the processing efficiency and resource utilization of computing devices for generation tasks. Therefore, it is difficult to reflect the essential improvement of the collaborative processing capabilities of terminals, servers, and networks from the perspective of "improvement of computer technology".

[0312] Therefore, it is necessary to provide a new system and program processing method that can: automatically convert simple user operations on the terminal into high-level control information and prompts suitable for generative artificial intelligence models; perform emotion analysis and generate soothing responses by combining text and symbol information in chat scenarios; and iteratively update voice parameters and prompts through user evaluation information, thereby forming an efficient and intelligent control and reasoning framework within the computer for virtual object generation and voice generation tasks, so as to improve generation quality, human-computer interaction experience, and computing resource utilization efficiency.

[0313] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.

[0314] In this invention, the server includes a component for acquiring attribute information and voice parameters related to speech generation from a user terminal, receiving the attribute information and voice parameters as input, and abstracting the information within a computing device to generate higher-level feature information representing user-desired features; a component for automatically generating prompt statements based on the higher-level feature information to instruct a generative artificial intelligence model to generate virtual objects; a component for parsing voice parameters into control information representing pitch, speech rate, speech manner, speech style, and emotional state, and generating prompt statements based on the control information to instruct the speech generation to use a generative artificial intelligence model to generate voice information with predetermined speech content and predetermined speech features; and a component for processing text acquired in conversational form. The system performs natural language processing on information and symbolic information to generate high-level emotional state information on the server side. Based on this emotional state information, it generates prompt statements to instruct the generative artificial intelligence model to generate response information that reassures the recipient. It also includes components for invoking the generative artificial intelligence model deployed on computing hardware to generate corresponding image, speech, and response information based on the prompt statements; and components for sending the generated speech and image information to the user terminal via a communication path, receiving evaluation information about the speech reproduction results, updating the control information and higher-level feature information internally on the server based on the evaluation information, thereby regenerating the optimized prompt statements and driving the generative artificial intelligence model for iterative generation. This allows for the construction of a unified prompt statement generation and control framework for generative artificial intelligence models between the server and the terminal, automatically mapping simple user input to high-level control information. This enables the computer system to perform virtual object generation, speech generation, and emotional response generation tasks with higher efficiency, higher accuracy, and stronger adaptability, thereby substantially improving the processing performance and user interaction experience of computers in the field of multimodal content generation.

[0315] A "system" refers to a collection of devices consisting of at least one server, at least one user terminal, and programs running on them, used to perform tasks such as generating prompts, calling generative artificial intelligence models, and presenting information over a network.

[0316] A "server" refers to an electronic computing device that provides computing and storage resources in a network environment. It is used to receive requests from user terminals and perform centralized processing such as generating prompts, parsing parameters, inferring models, and returning results.

[0317] "User terminal" refers to an information processing device that is operated by a user and interacts with a server, including but not limited to mobile terminals, fixed terminals or other computing devices with interface display and input functions.

[0318] "Attribute information" refers to a collection of high-level characteristic information related to users or virtual objects, including but not limited to age category, gender category, personality category, interest category, usage scenario category, etc., which are used to guide generative artificial intelligence models to generate content that conforms to the set characteristics.

[0319] "Speech parameters" refers to a set of settings used to control the characteristics of speech generation, including but not limited to pitch, speech rate, speech manner, speech style, emotional state, and accent type.

[0320] "Higher-level feature information" refers to a higher-level feature representation generated internally by the server through abstract processing based on raw inputs such as attribute information and voice parameters. It is used to summarize user intent and target style so as to serve as control conditions for generative artificial intelligence models.

[0321] "Control information" refers to numerical or categorical information derived from data such as speech parameters that can be directly used to control the behavior of generative artificial intelligence models, including but not limited to pitch scaling factor, speech rate scaling factor, vocalization type, style type, and emotion type identifier.

[0322] "Virtual objects" refer to virtual characters or entities generated through generative artificial intelligence models, which are represented as images, videos, or three-dimensional forms in a digital environment and are used for display and interaction on an interface.

[0323] "Prompt statements" refer to instruction information generated by the server or user terminal based on attribute information, voice parameters, higher-level feature information, etc., which describes the generation requirements in natural language or structured text and is used to instruct generative artificial intelligence models on target content and target features.

[0324] "Generative AI models" refer to AI models that can automatically generate content data based on input prompts and control information, including but not limited to models used for image generation, speech generation, text generation, or multimodal generation.

[0325] "Generative AI models for speech generation" refers to generative AI models specifically designed to generate speech information based on spoken content, speech parameters, and control information, including but not limited to text-to-speech conversion models and their accompanying vocoder models.

[0326] "Generative AI models for image generation" refers to generative AI models specifically designed to generate image information based on attribute information, higher-level feature information, and relevant prompts, including but not limited to generative adversarial network models and diffusion models.

[0327] "Text information" refers to text data entered or received in character form during chat or interaction, including natural language sentences, phrases, words, etc., used for content understanding and sentiment analysis.

[0328] "Symbolic information" refers to non-textual symbolic data contained in chat or interaction, including but not limited to emoticons, graphic symbols, punctuation patterns, and other marks that can reflect the user's emotions or tone.

[0329] "Emotional state information" refers to high-level characteristic information obtained from textual and symbolic information through natural language processing and sentiment analysis, used to represent a user's current emotions or moods, including but not limited to categories such as happiness, sadness, tension, peace of mind, and anger.

[0330] "Response information" refers to content data automatically generated by a generative artificial intelligence model to reply to users based on emotional state information and prompts, including but not limited to reassuring text, explanatory text, or guiding text.

[0331] "Image information" refers to static or dynamic visual data that can be displayed on a user terminal, output by a generative artificial intelligence model for image generation, including but not limited to bitmap data, vector data, or texture data.

[0332] “Speech information” refers to digital audio data that can be played on a user terminal, output by a generative artificial intelligence model for speech generation, including but not limited to raw audio waveforms, encoded audio streams, or audio files.

[0333] "The prompted subject" refers to the user or viewer who receives and perceives image information, voice information, or response information through the user terminal, and is the object of perception of the system's output content.

[0334] "The terminal being prompted" refers to a user terminal or display / playback device that presents image information, voice information, or response information to the subject being prompted.

[0335] "Evaluation information" refers to the feedback data given by the prompted subject after viewing or listening to the generated content regarding the quality of the content, the degree of style matching, etc., including but not limited to ratings, tag selections, free text opinions, or implicit interaction behavior data.

[0336] "Communication path" refers to the network connection between the server and the user terminal used to transmit attribute information, voice parameters, prompts, image information, voice information, and evaluation information, including but not limited to wired network connections and wireless network connections.

[0337] A “computing device” refers to an electronic device with a processor, memory, and input / output interfaces, used to execute program instructions to perform functions such as attribute abstraction, prompt generation, model reasoning, and result return.

[0338] In one embodiment of the invention, the server can be constructed in hardware as an electronic device with a multi-core processor and a graphics processing unit, such as a computer system equipped with a general-purpose central processing unit (CPU) and a graphics processing unit (GPU). The server can run a general-purpose operating system, such as a Unix-like operating system, and install deep learning framework software, such as a framework based on tensor computation, on it. The server maintains parameter files of multiple generative artificial intelligence models in a storage device and loads these parameters into memory and video memory for performing inference processing such as virtual object generation, speech generation, and text generation.

[0339] In one embodiment of the present invention, the terminal can be configured in hardware as a portable information processing device, such as a smart terminal device or a tablet device, equipped with a display device, a touch input device, a speaker, and a wireless communication module. The terminal runs a general-purpose mobile operating system and implements functions such as user interface rendering, user input acquisition, network request sending, and audio playback through applications. The user operates the terminal through its display screen and touch screen to provide attribute information and voice parameters to the server and receive image information, voice information, and response information returned by the server.

[0340] In one embodiment of the present invention, the server maintains multiple program modules in a storage device. These program modules are loaded into memory and executed by the processor to implement functions such as an attribute abstraction module, a prompt statement generation module, a parameter parsing module, a sentiment analysis module, a model reasoning module, and a feedback learning module. Logically, the server divides these modules into several components, each of which performs specific data processing and data operations.

[0341] In one embodiment of the invention, the attribute information and voice parameters provided by the terminal to the server are managed in the form of structured data. The terminal uses key-value mapping or a similar structure in its local memory to store the options selected by the user in the interface controls, and converts this data into an abstract description composed of strings and category identifiers before sending it to the server. The user can select, for example, "voice pitch: high," "speech speed: fast," "accent: Kansai dialect," and input the text to be read aloud, such as "The weather is really nice today, let's go for a walk." After the terminal transmits the above information to the server, the server converts it into an internally unified control information representation.

[0342] In one embodiment of the present invention, the server converts attribute information and speech parameters received from the terminal into higher-level feature information through an attribute abstraction module. In this module, the server uses a feature mapping table and an embedding vector generation algorithm to map discrete category information into a continuous feature space. For example, the server assigns a pitch coefficient (e.g., 1.3) to "high pitch," a speech rate coefficient (e.g., 1.2) to "fast speech," and a dialect category identifier to "Kansai dialect." The server combines these scalars into a higher-level feature vector through matrix operations, enabling the generative artificial intelligence model to interpolate and combine user intentions in a continuous feature space, thereby improving the granularity and control precision of the generated results.

[0343] In one embodiment of the present invention, the server automatically generates prompt statements for driving the generative artificial intelligence model based on higher-level feature information through a prompt statement generation module. In this module, the server uses a combination of template matching and rule combination to convert control information into instructions in natural language form. Specifically, the server may use the following prompt statement examples: "Please read the following sentences aloud in a high-pitched, fast-paced tone, using the Kansai dialect." The server can generate different prompts based on different combinations of parameters, for example: Please read the following sentences aloud in a medium pitch, at a slightly slower pace, with a gentle tone, and in standard Mandarin. In this way, the server continuously maintains a set of mapping rules internally, which bidirectionally associates numerical control information with natural language prompts. This allows the generative artificial intelligence model to obtain style control information through prompts when receiving text input, and the server can also deduce the corresponding control parameters from the prompts for recording and optimization when needed.

[0344] In one embodiment of the present invention, the server employs a sequence-to-sequence neural network model for speech generation using a generative artificial intelligence model. This model can consist of a text encoder, a temporal decoder, and an acoustic feature prediction network. When loading the model, the server first converts the text to be read into a discrete symbol sequence, which is then input into the text encoder. The text encoder can be composed of a multi-layer feedforward neural network, a recurrent neural network, or a multi-head attention network. The server inputs the prompt and control vector into the model together, and incorporates the control vector as a conditional embedding into the decoding process within the network. This allows the model to adjust the generation trajectory based on pitch coefficients, speech rate coefficients, emotion categories, etc., when outputting Mel spectrum or other acoustic features.

[0345] In this implementation, the server feeds acoustic features into the vocoder model. The server can use a convolutional neural network-based vocoder architecture, which progressively converts the Mel spectrum into a time-domain waveform through multiple layers of convolution and residual connections. During inference, the server uses a graphics processing unit to perform numerous convolution calculations and matrix multiplication operations, enabling rapid computation for each sample point. Since control information is embedded in the acoustic feature generation stage, the vocoder can directly generate speech waveforms conforming to the target style without additional post-processing, thereby reducing processing steps and improving the overall speech generation speed of the system.

[0346] In one embodiment of the invention, the server can employ a neural network model based on a generative adversarial structure or a diffusion structure for image generation using a generative artificial intelligence model. The server takes a vector representing user attribute information and higher-level feature information as conditional input, concatenates it with a random noise vector, and then inputs it into the image generation network. The image generation network gradually generates images with target appearance features through multi-layer upsampling convolutions and normalization operations. For example, the server can generate virtual object images with corresponding hairstyles, expressions, and clothing based on abstract labels such as "adult woman," "gentle," "virtual character," and "stage scene." By performing conditional control in the feature space, the server can flexibly adjust the generation results while maintaining the stability of the network structure, thereby improving the diversity and consistency of image generation.

[0347] In one embodiment of the invention, the server processes text and symbol information received in chat format through a sentiment analysis module. The server first segments and encodes the text information, and encodes the symbol information (e.g., emoticons, repeated punctuation, stretched characters, etc.) using features. These features are then input into a multi-layer neural network for sentiment classification. The server pre-trains this network using a cross-entropy loss function and a gradient descent algorithm, enabling the network to output probability distributions of emotional states such as "happy," "sad," "nervous," and "reassuring." During the inference phase, the server determines the user's current emotional state based on the category with the highest probability and provides this emotional state as high-level sentiment state information to the prompt generation module.

[0348] Based on this sentiment analysis, the server generates reassuring response prompts. For example, when the server detects that a user is in a state of tension, it can generate the following prompt: "Please generate a text in a gentle and reassuring tone to express understanding and support to the user." The server then inputs the prompt into a generative AI model for text generation, thus generating a specific response. The server then sends the resulting response text to the terminal via the network, where it is displayed as a dialogue on the interface. This structure allows the system to automatically complete a series of technical processes, from emotion recognition to the generation of reassuring text, within the computer without human intervention.

[0349] In one embodiment of the invention, the server updates parameters using user feedback on the speech reproduction results through a feedback learning module. After listening to the generated speech information on the terminal, the user can select "satisfied" or "unsatisfied" through a simple interface or give a rating on a slider. The terminal sends this evaluation information to the server. The server adjusts the mapping relationship between control information and higher-level feature information based on the evaluation results in the feedback learning module. For example, when a large number of users mark the generated result as "too fast" under the combination of "high pitch, fast speech rate, and Kansai dialect," the server can automatically reduce the speech rate coefficient of the corresponding combination from 1.2 to 1.1 and use the new value in subsequent calls. This adjustment based on statistical feedback is implemented internally by updating the parameter table and using a simple optimization algorithm, enabling fine-grained fine-tuning of the generation behavior without retraining a large model, thereby improving the consistency between the generated results and user expectations.

[0350] In one embodiment of the invention, the server optimizes the performance of the generation task through modular design and a caching mechanism. Upon receiving the prompt and text content, the server first checks the cache module to see if there are any generation results with the same or similar control information combined with the text. If the result exists in the cache, the cached audio byte data or image data can be directly returned to the terminal, reducing redundant computation. For requests that do not hit the cache, the server invokes the generative artificial intelligence model for inference. In high-concurrency scenarios, this mechanism can significantly reduce GPU inference load, improve response speed, and reduce network traffic and memory access frequency, thereby achieving communication load reduction and computational efficiency improvement at the computer technology level.

[0351] In another embodiment of the invention, the server can share some intermediate features across different generation tasks. For example, when a user requests the same text to be read aloud multiple times in the same session with different speech styles, the server can encode the text once and store the encoding result in a cache. In subsequent requests, the server only needs to re-execute acoustic feature generation and vocoder inference for different control information, without repeating the text encoding, thereby reducing redundant computation. Since text encoding is usually time-consuming in the model structure, this sharing strategy can substantially shorten the overall generation time and improve the utilization of server resources.

[0352] In one embodiment of the present invention, the data format and protocol between the server and the terminal can be modified according to implementation requirements. The server can represent audio data using a binary encoding format or a compression encoding method to reduce network bandwidth consumption. After receiving the voice information returned by the server, the terminal directly calls the local audio decoding and playback module to output audio through the speaker. In actual use, users can directly listen to the generated voice in scenarios such as virtual customer service, virtual host, educational tutoring, or game character voice-over. Through the system structure of the present invention, the actual acoustic output is generated by the terminal's audio hardware, thereby connecting the abstract generation processing inside the server with the control of audio playback devices in the real world.

[0353] In another embodiment of the invention, the server can select different generative artificial intelligence model architectures based on varying hardware conditions. For example, in scenarios with ample graphics processing unit resources, the server can use a diffusion-based generative model with a larger parameter scale to achieve higher image quality; in scenarios with limited resources, a lightweight network structure with fewer parameters can be adopted to balance speed and resource consumption. During runtime, the server can dynamically select the model version and batch processing strategy based on the task queue length and current resource utilization to achieve computational load balancing and latency control. These technical measures directly improve the performance of the computer system in a multi-tasking environment.

[0354] In all the above implementations, the interaction between the server, terminal, and user not only achieves the convenience of the human-computer interface, but also establishes an efficient control and reasoning framework for generative artificial intelligence models within the computer through specific data structures, neural network structures, and feedback learning mechanisms. This framework, through techniques such as dual encoding of prompts and control information, parameter abstraction and iterative optimization, modular reasoning, and caching mechanisms, achieves fine-grained control of generation quality, substantial improvement in generation speed, and significant improvement in the efficiency of computing resource utilization, thus demonstrating the technological improvements over traditional manual settings and simple automation schemes.

[0355] use Figure 13 The processing procedure is explained.

[0356] Step 1: Users configure generation requirements on the terminal. Users select or input parameters and text content related to speech generation on the terminal's graphical interface.

[0357] Input: The user selects attribute information (such as role type, personality type, scene type) and voice parameters (such as voice pitch "high", speech speed "fast", accent "Kansai dialect") on the terminal interface and the text to be read aloud (such as "The weather is really nice today, let's go for a walk.").

[0358] The terminal obtains touch events and text input through the input interface provided by the operating system, and reads the current value of each control into a data structure in memory (such as a key-value mapping).

[0359] The terminal performs basic validation on these raw inputs (e.g., checking if the text is empty or if the parameters are within the allowed set of values), and then forms a structured request object locally as the data basis for subsequent transmission to the server.

[0360] Output: A local structured data object containing attribute information, voice parameters, and text content.

[0361] Step 2: The terminal sends a structured request to the server. The terminal constructs a network request based on the structured data object generated in step 1.

[0362] Input: A structured data object containing attribute information, voice parameters, and text content.

[0363] The terminal calls the network communication library to serialize the structured data into a transmission format (e.g., encoded as a string) and construct a request message containing the target server address, request path, and header fields.

[0364] The terminal performs data fragmentation and encapsulation with the network stack through the wireless communication module, and sends the request message to the server through the communication path.

[0365] Output: A request message transmitted over the network to the server, which includes the user's attribute information, voice parameters, and the text to be read aloud.

[0366] Step 3: The server receives and parses the user request. The server receives requests from the terminal at the network interface and hands them over to the application for processing.

[0367] Input: A request message sent by the terminal (containing encoded data of attribute information, voice parameters, and text content).

[0368] The server uses a parsing module to decode the request message, extract the message body, and deserialize the string data into an internal data structure (such as an object with fields including attribute information, voice parameters, and text content).

[0369] The server performs a validity check on the parsed results, such as checking whether required fields exist and whether parameter types are correct, and then distributes the parsed data to the attribute abstraction module and the parameter parsing module.

[0370] Output: User attribute information object, voice parameter object, and text content string represented in server memory.

[0371] Step 4: The server abstracts attribute information and voice parameters into higher-level feature information. The server uses the attribute abstraction module to perform feature mapping on user attribute information and voice parameters.

[0372] Input: User attribute information object and voice parameter object (e.g., voice pitch = high, speech rate = fast, accent = Kansai dialect, role type, etc.).

[0373] Based on pre-defined mapping rules and feature tables, the server converts categorical parameters into numerical control quantities. For example, it maps "high pitch" to a pitch scaling factor, "fast speech rate" to a speech rate scaling factor, and "Kansai dialect" to a dialect category index.

[0374] The server combines multiple control variables into a single higher-level feature vector through vector concatenation and linear transformation. Internally, it performs matrix multiplication and activation function operations to obtain a continuous feature representation that can be used by generative artificial intelligence models.

[0375] Output: A vector of higher-level feature information representing the user-generated requirements, along with the corresponding set of control information.

[0376] Step 5: The server generates prompts for virtual objects and speech generation. The server generates prompts in natural language based on higher-level feature information and control information through a prompt generation module.

[0377] Input: upper-level feature information vector, control information set, and text to be read aloud.

[0378] The server uses a template library and rule set to transform control information into human-readable descriptive fragments, then performs string concatenation and placeholder replacement to assemble complete prompt statements. For example, the server generates prompt statements based on control information such as "high-pitched voice, fast speech, and Kansai dialect": "Please read the following sentences aloud in a high-pitched, fast-paced tone, using the Kansai dialect." The server can also generate prompts for virtual object image generation based on attribute information, for example: "Please generate a virtual character with a gentle temperament, suitable for conversing with users in a relaxed setting." Output: Prompt text used to drive generative artificial intelligence models to generate speech and image information.

[0379] Step 6: The server performs sentiment analysis on chat text and symbolic information. In a conversational scenario, the server receives chat text and symbol data from the terminal and determines the emotional state.

[0380] Input: Text information (sentences, phrases, etc.) and symbol information (emoticons, punctuation patterns, stretched characters, etc.) sent by the user through the terminal.

[0381] The server first performs word segmentation and encoding, splitting the text into words or sub-word units and mapping them to a sequence of integer IDs. At the same time, it performs feature encoding on symbol information (for example, mapping different emoticons to discrete indices and using the number of repeated symbols as numerical features).

[0382] The server inputs the above encoded sequence into a pre-trained emotion classification neural network. The network performs matrix multiplication, nonlinear activation, and normalization operations through embedding layers, sequence modeling layers, and classification layers, and outputs probability distributions corresponding to various emotion categories.

[0383] The server determines the emotional state (e.g., "nervous", "sad" etc.) based on the highest probability or a set threshold, and generates emotional state information representing that emotional category.

[0384] Output: High-level emotional state information representing the user's current mood.

[0385] Step 7: The server generates reassuring response prompts based on the emotional state. The server generates prompts for emotional responses based on emotional state information, and drives a generative artificial intelligence model to generate response text.

[0386] Input: High-level emotional state information (e.g., "nervous" or "anxious") and contextual dialogue information.

[0387] The server uses rule mapping to associate different emotional states with different coping strategies and generates prompts describing the target tone and style. For example, when the emotional state is "nervous," the server generates the following prompt: "Please generate a text in a gentle and reassuring tone to express understanding and support to the user." The server inputs the prompt statement and the session context text into a generative artificial intelligence model for text generation. The model performs encoding and decoding operations on the server to generate the specific response text content.

[0388] Output: A reassuring or empathetic response text message.

[0389] Step 8: The server calls a speech generation function to use a generative artificial intelligence model to generate speech information. The server uses a model inference module to input prompts, text to be read aloud, and control information into a generative artificial intelligence model for speech generation, thereby generating speech information.

[0390] Input: Voice-generated prompts (e.g., "Please read the following sentence aloud in a high-pitched, fast-paced tone, speaking in the Kansai dialect."), the text to be read, and control information vectors.

[0391] The server first encodes the text to be read into a sequence of symbols and inputs it into the text encoding subnetwork, where embedding, attention calculation, or loop operation is performed to generate a text representation. At the same time, the server injects the control information vector as a conditional input into the decoding subnetwork.

[0392] The server generates intermediate acoustic features (such as a Mel spectrum matrix) through a decoding subnetwork, and then inputs these features into a vocoder network to generate a time-domain speech waveform array through multi-layer convolution and upsampling operations.

[0393] Output: Digital speech information data (e.g., an array of PCM samples with a certain sampling rate).

[0394] Step 9: The server invokes an image generation function to generate virtual object images using a generative artificial intelligence model. The server generates virtual object image information based on higher-level feature information and attribute information.

[0395] Input: Virtual object related attribute information, higher-level feature information vector, and image generation prompt statement (e.g., "Please generate a virtual character image with a gentle temperament, suitable for talking to the user in a relaxed scene.").

[0396] The server concatenates the high-level feature information with a random noise vector and inputs it into the image generation network. Through multi-layer upsampling convolution, normalization, and non-linear activation operations, it gradually generates the image feature map of the target resolution and the final pixel matrix.

[0397] The server encodes or compresses the generated image matrix, preparing it for transmission as an image file or image stream.

[0398] Output: Image information data representing the appearance characteristics of the virtual object.

[0399] Step 10: The server encapsulates the voice and image information and sends it to the terminal. After completing the speech and image generation, the server returns the results to the terminal through the communication module.

[0400] Input: Voice information data and image information data, as well as response text information (if any).

[0401] The server sets the corresponding transmission format according to different data types, encapsulates voice data into audio streams or audio file formats, encapsulates image data into image formats, preserves text as strings, and organizes these data by field in the response message.

[0402] The server sends the response message to the terminal through the network protocol stack, and internally performs operations such as data fragmentation, reassembly, and error detection to ensure the integrity of the data during transmission.

[0403] Output: A response message containing voice information, image information, and response text.

[0404] Step 11: The terminal receives the server's response and performs data parsing and caching. The terminal receives the response message returned by the server from the communication module and parses the data in it.

[0405] Input: The response message received from the server, which contains encoded voice, image, and text information.

[0406] The terminal uses a parsing module to decode the message, extract audio data, image data, and text data according to predetermined fields, store the audio byte stream in the local cache, load the image data into the image buffer, and save the text in memory for display.

[0407] Terminals can choose to create temporary files in the local file system so that they do not have to request the server again when repeating audio or displaying images, thereby reducing network access and improving response speed.

[0408] Output: Audio buffer, image buffer, and text data that can be played and displayed on the terminal.

[0409] Step 12: The terminal presents the generated results to the user and collects evaluation information. The terminal uses local multimedia and interface modules to present server-generated content to the user and receive user feedback.

[0410] Input: Audio cache, image cache, and text data obtained in step 11.

[0411] The terminal displays a virtual object image and response text on the screen, while simultaneously calling the audio playback component to read voice data from the cache and play it through the speaker, allowing the user to see and hear the generated results.

[0412] Users rate the voice quality on the terminal interface using buttons or rating controls. The terminal converts the user's evaluation into structured evaluation information (such as rating values ​​or satisfaction / dissatisfaction indicators) and prepares to send it to the server for feedback and learning.

[0413] Output: The multimodal generation results and evaluation information data generated by user operations are displayed on the terminal.

[0414] Step 13: The server updates control information and prompt statement generation rules based on the evaluation information. The server uses user feedback information to iteratively optimize the generation process of control information and prompt statements.

[0415] Input: Evaluation information data sent from the terminal (e.g., user satisfaction with the combination of "high-pitched voice, fast speech speed, and Kansai dialect").

[0416] In the feedback learning module, the server adjusts the control coefficients or mapping rules stored in the parameter table based on the evaluation values. For example, it may appropriately reduce the speech rate coefficient or adjust the weight of the emotion label for combinations with lower evaluations.

[0417] The server uses statistical analysis of multiple evaluation results and simple optimization algorithms (such as moving average or weighted update formula) to correct the mapping relationship between control information and higher-level feature information, thereby using the updated parameters in subsequent prompt generation and model inference.

[0418] The server also adjusts the prompt statement generation rules based on the updated control information. For example, under the same parameter combination, the prompt statement is changed from "speech speed fast" to "speech speed slightly faster" to more accurately reflect the new control objectives.

[0419] Output: The updated control information mapping table and prompt statement generation rules are used for generation processing in future requests, so that the system gradually approaches the generation behavior expected by the user.

[0420] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0421] The following technical problems exist in existing interactive systems based on computing devices: First, servers typically drive content generation algorithms directly with fixed fields, lacking abstraction of user attribute information and modeling of higher-level concepts. This makes it difficult to reuse the generated virtual avatars and interactive content across different scenarios, resulting in low model calling efficiency and limited system ability to express complex user preferences, thus restricting the utilization efficiency and scalability of computing resources.

[0422] Second, when processing user voice output requests, most systems only provide static voice templates or simple parameter mappings. They do not systematically abstract user selection information into higher-level voice parameters at the program level, nor do they drive generative artificial intelligence models to generate voice expressions through unified prompt statements. This results in a high degree of coupling between the voice synthesis module and the upper-level business logic, making it difficult to share across applications and scenarios, and leading to poor flexibility and maintainability of voice generation.

[0423] Third, when processing text information and visual symbols entered by users through chat, the common practice of servers is to only perform sentiment polarity analysis on the text. They cannot transform text and visual symbols such as emoticons into higher-level linguistic information and sentiment indicators in a unified natural language processing flow, nor can they quantify and classify emotions in multiple dimensions. Therefore, it is difficult to accurately capture the fine-grained emotional state of users, especially complex emotions that contain multiple emotions, which are difficult to accurately identify and utilize.

[0424] Fourth, existing systems generally lack a mechanism for structurally modeling the computational path of "emotional state → prompt statement → output of generative artificial intelligence model". Servers usually generate responses with hard-coded response templates or simple rules, making it difficult to dynamically construct prompt statements with high semantic information based on numerical and categorized emotional states, and to use generative artificial intelligence models to generate personalized response and support information. This results in significant deficiencies in dialogue quality, contextual coherence, and user reassurance.

[0425] Fifth, in terms of content recommendation and multimodal presentation, traditional systems mostly rely on behavioral statistics for recommendations, failing to deeply couple real-time emotional states with the external content selection process. Furthermore, they lack a unified technical architecture that automatically generates recommendation reason text through generative artificial intelligence models and outputs it in a composite form along with the recommended content. This limits the expressive power of personalized recommendation logic at the computational level and makes it difficult to achieve interpretable and scalable management of the recommendation process.

[0426] Sixth, for user input containing multiple emotions simultaneously, existing sentiment analysis modules typically only output a single main sentiment tag. The server lacks a mechanism in the programming logic to extract multiple emotions simultaneously and convert them into structured prompts for use by generative artificial intelligence models. It is unable to combine and map multiple emotions and generate corresponding comprehensive response information, resulting in weak emotion management capabilities in multi-turn dialogues and affecting the overall interactive experience.

[0427] In summary, a computational technology solution is needed to uniformly abstract, quantify, and structurally model user attribute information, voice parameters, chat content, and emotional state on the server side. This solution should drive generative artificial intelligence models by generating prompt statements, thereby systematically improving the processing flow of multimodal content generation and emotion-driven interaction at the program level, and enhancing the system's scalability, reusability, and computational resource utilization efficiency.

[0428] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.

[0429] In this invention, the server includes a processing device for acquiring user attribute information from a user terminal and converting the attribute information into abstract attribute information based on a higher-level concept. The processing device is configured to generate prompt statements based on the abstract attribute information to enable a generative artificial intelligence model to generate a visual representation of a virtual entity. A mapping device is also included for acquiring selection information related to voice output from the user terminal and mapping the selection information to voice parameters of a higher-level concept. The mapping device is configured to generate prompt statements based on the voice parameters and text information to enable the voice generation model to generate synthesized speech and instruct the generation of the synthesized speech. Finally, a conversion device is included for receiving text information and visual symbols from the user terminal in a chat-like format and converting the text information and visual symbols into language information and sentiment indicators of a higher-level concept. The conversion device is configured to quantify and classify the user's emotional state using natural language processing and sentiment analysis techniques, and to convert the quantified and classified sentiment information into... The system includes: a sensory state and dialogue history information generation tool for generating prompt statements that provide reassurance and support information to the user through a generative artificial intelligence model; an output device for converting the response and support information output from the generative artificial intelligence model into a composite presentation form of higher-level concepts including text and voice information and sending it to the user terminal for presentation; a content recommendation device for obtaining content of higher-level concepts from an external information providing device or information recording device based on the sensory state, selecting the content, and generating prompt statements that enable the generative artificial intelligence model to generate recommendation reasons for the content and presenting the content and recommendation reasons to the user terminal; and a multi-emotion processing device for simultaneously extracting multiple emotions using natural language processing technology and generating prompt statements that enable the generative artificial intelligence model to generate comprehensive response information composed of response elements corresponding to each emotion when the sensory state contains multiple emotions. This allows for the formation of a unified computing chain within the server, encompassing "abstracted data → prompt generation → generative AI model invocation → multimodal output." This enables structured processing of user attributes, voice requirements, and complex emotional states, improving the decoupling and reusability between multiple modules such as virtual entity generation, speech synthesis, emotional response, and content recommendation. Consequently, it enhances the utilization efficiency of computing resources, reduces program modification costs during system updates and expansions, and significantly improves the overall technical performance of computing-based interactive systems in multimodal emotion-driven services.

[0430] "User terminal" refers to an information processing device operated by a user and used to interact with the system, including but not limited to portable computing devices, fixed computing devices, and any electronic device capable of sending and receiving data with a server through a communication network.

[0431] A "server" refers to a collection of information processing devices configured with processing resources and exchanging data with user terminals through a communication network. It is used to perform functions such as data reception, data processing, model invocation, and result return. It can consist of one or more physical or virtual computing nodes.

[0432] "User attribute information" refers to a higher-level concept that describes various types of information used to represent the characteristics of a user or a virtual object associated with a user, including but not limited to appearance characteristics, behavioral characteristics, preference characteristics, and other parameter information used to generate virtual entities.

[0433] "Abstracted attribute information" refers to attribute data represented in the form of higher-level concepts, obtained through feature induction, hierarchical expression, or category mapping based on user attribute information. It is used to reduce detail differences and enhance the universality and reusability in different models or scenarios.

[0434] "Virtual entity" refers to the broader concept of virtual objects that exist in the form of data in electronic or digital environments and can be presented to users through visual, auditory, or other means, including but not limited to virtual characters, virtual persons, virtual agents, and other digital entities that can interact with users.

[0435] "Visual representation" refers to a higher-level concept used to represent the visual characteristics of virtual entities, including but not limited to static images, dynamic images, two-dimensional or three-dimensional graphics, and related rendering data.

[0436] "Generative artificial intelligence model" refers to the broader concept of an artificial intelligence model that is trained in a data-driven manner to automatically generate text, images, audio, or other data outputs based on input conditions or prompts, including but not limited to text generation models, image generation models, and speech generation models.

[0437] "Prompt statements" are higher-level concepts that refer to the instructive or guiding information input into a generative artificial intelligence model, including but not limited to natural language text, structured text, or parameterized descriptions, used to instruct the generative artificial intelligence model to perform specific generative tasks and constrain its output characteristics.

[0438] "Voice output related selection information" refers to the overarching concept of information corresponding to various selection operations made by users in the user terminal regarding voice performance, including but not limited to language selection, timbre selection, speech rate selection, gender preference selection, and other parameter selection information related to voice performance.

[0439] "Speech parameters" refers to a higher-level concept that is a parameterized representation obtained by mapping speech output-related selection information and can be directly used in the speech generation or speech synthesis process, including but not limited to pitch parameters, speech rate parameters, volume parameters, vocal timbre category parameters, and language category parameters.

[0440] "Generative AI model for speech generation" refers to the overarching concept of a generative AI model that uses speech or audio data as its primary output, and is used to generate synthesized speech or audio signals based on input text and speech parameters.

[0441] "Synthetic speech" refers to the broader concept of audio data generated based on text and parameters using generative artificial intelligence models or other speech synthesis technologies to replace the voices of natural people.

[0442] "Text information" refers to the broader concept of input information consisting of character sequences used to express semantic content, including but not limited to natural language text, punctuation marks, and text data contained in chat messages or commands.

[0443] "Visual symbols" refers to a broader concept of information that expresses emotion, tone, or intention in the form of graphics, icons, emoticons, or other visual elements, including but not limited to emoticons, graphic symbols, and icons used to assist in the expression of gestures.

[0444] "Language information" refers to the overarching concept that is formed by parsing, normalizing, or semantically understanding textual information, and is used to represent the intermediate representation of language content and structure. It includes, but is not limited to, word sequences, syntactic structure information, and semantic tag information.

[0445] "Emotional metrics" refers to a higher-level concept that uses symbolic or numerical information to quantitatively or qualitatively represent a user's emotional state, including but not limited to emotional category labels, emotional intensity values, polarity scores, and multidimensional emotional vectors.

[0446] "Natural Language Processing technology" refers to the overarching concept of algorithms and models used to analyze, understand, generate, and transform natural language data, including but not limited to word segmentation, part-of-speech tagging, syntactic analysis, semantic analysis, and text generation technology.

[0447] "Sentiment analysis technology" refers to the broader concept of algorithms and models used to identify, classify, and quantify emotional states from linguistic information, visual symbols, or other input data, including but not limited to emotion classification models, emotion polarity analysis models, and multidimensional emotion recognition models.

[0448] "Emotional state" refers to a higher-level concept that represents a user's emotional state within a specific time period, obtained by processing user input data through sentiment analysis technology. It includes, but is not limited to, single emotion categories, complex emotion combinations, and corresponding numerical indicators.

[0449] "Conversation history information" refers to a higher-level concept of a collection of data related to the interaction between the user and the system that is accumulated and recorded during one or more conversations, including but not limited to historical message text, timestamps, context labels, and associated sentiment state records.

[0450] "Response information" refers to the higher-level concept of output content generated by the system based on user input, emotional state, and dialogue history information, used to directly respond to the user's current input, including but not limited to text replies, voice replies, and other forms of response data.

[0451] "Support information" refers to a higher-level concept of output content that the system actively generates based on the user's status or context, in addition to direct response information, to provide assistance, advice, reassurance, or guidance.

[0452] "Composite presentation format" refers to a higher-level concept that combines multiple types of information and presents them in a unified interface or unified output signal, including but not limited to the combination of text and voice, the combination of images and text, and the synchronous presentation of multimedia content.

[0453] "External information providing device" refers to a higher-level concept of a device or system located outside the system that can provide data or services to the server through a communication interface, including but not limited to external database systems, content distribution systems, and third-party service platforms.

[0454] "Information recording device" refers to a broader concept of a storage system used to store data related to users or content that can be accessed by a server, including but not limited to local storage systems, network storage systems, and distributed storage clusters.

[0455] "Content" is a broader concept referring to various information objects that users can access or consume, including but not limited to text content, audio content, video content, image content, and interactive resources.

[0456] "Reasons for recommendation" is a higher-level concept that explains or describes why a certain content was selected as a recommended item, and is used to improve the understandability and interpretability of the recommendation results.

[0457] "Multiple emotions" refers to the overarching concept of two or more different categories of emotional components coexisting at the same time or in the same input, such as simultaneously including pleasure and fatigue, anticipation and anxiety, etc.

[0458] "Response elements" refers to the overarching concept of basic response fragments or logical units corresponding to specific emotional categories or dialogue needs when generating response information, including but not limited to comforting statements, suggestive statements, and confirmatory statements.

[0459] "Comprehensive response information" refers to the overarching concept of a holistic output response formed by combining multiple response elements generated separately for various emotions according to predetermined rules or dynamic strategies. It is used to simultaneously address and take into account multiple emotional components in a single response.

[0460] In various embodiments of the present invention, the system mainly consists of a server, a terminal, and a communication network. The server, as the core computing node, is equipped with one or more processors, a memory, and acceleration hardware (such as a graphics processing unit). The terminal, as a user interaction device, is equipped with a display unit, an input unit, an audio output unit, and a communication interface. The server executes programs stored on a storage medium to achieve coordinated invocation of the generative artificial intelligence model, the natural language processing module, and the sentiment analysis module, thereby realizing the various functions of the present invention.

[0461] I. System Overall Structure In one embodiment, the server includes: an attribute processing module, a prompt statement generation module, an image generation interface module, a speech generation interface module, a natural language processing module, a sentiment analysis module, a multi-sentiment processing module, a content recommendation module, a multimodal output module, and a data management module. In another embodiment, the terminal includes: a user interface module, an input acquisition module, a display module, an audio playback module, and a communication module.

[0462] Users select virtual entity attributes and voice preferences through the user interface module on the terminal, and input text and visual symbols. The terminal sends the input as structured data to the server through the communication module. After receiving the data, the server stores and manages it in the form of key-value pairs, vectors, and tags in its internal data management module.

[0463] II. Attribute Abstraction and Generation of Visual Representation of Virtual Entities In one implementation, the server uses an attribute processing module to abstract user attribute information received from the terminal. Specifically, the server stores an attribute ontology table in memory, which consists of multi-layered mapping rules. This attribute ontology table maps specific attributes (such as "blue hair," "green eyes," and "athletic body type") to higher-level conceptual attributes (such as "cool-toned hair color," "bright eye features," and "low-fat body type"). The attribute ontology table can be stored using a hierarchical label structure or a graph structure.

[0464] The server uses an attribute processing module to convert raw attribute information into abstract attribute information. Then, the server uses a prompt generation module to generate prompts based on the abstract attribute information to drive the generative artificial intelligence model. For example, the server generates the following prompts: "A blonde, blue-eyed, adventurer-style anime-style virtual character, full-body portrait, high-quality illustration style." In one implementation, the server invokes an image-based generative artificial intelligence model through an image generation interface module. In a specific implementation, this model can be a generative artificial intelligence model based on a diffusion model structure, internally employing a U-Net network, a variational autoencoder submodule, and a noise scheduler. The server inputs the aforementioned prompt as text conditions into the text encoding subnetwork of the diffusion model. Internally, the text encoding subnetwork encodes the prompt into a semantic vector. Subsequently, the diffusion model generates latent image features in the latent space through a back-diffusion process, and the decoder converts these latent features into a pixel matrix, thereby generating a visual representation of the virtual entity.

[0465] The server stores the generated image data in the image storage area and sends the corresponding reference information to the terminal. The terminal then displays the virtual image on the display screen via the display module, allowing the user to visually confirm the generated result.

[0466] By combining the aforementioned attribute abstraction with prompt generation, the server can share the same set of abstract attributes and prompt templates among different users. This reduces the number of repeated model calls, improves the efficiency of computing resource utilization, and makes it possible to reuse the same generation process in different application scenarios (such as entertainment and education environments). This abstraction process improves upon the traditional field-by-field image generation method from the perspective of internal computer data structures and model call processes. It achieves a structured mapping between high-level semantics and model input, reducing system coupling.

[0467] III. Speech Parameter Mapping and Speech Generation In one implementation, the server uses a speech generation interface module to process user speech output selection information. Users can select gender, speech rate, tone style, and language on the terminal. The terminal then sends these selections as structured data to the server.

[0468] The server maintains a speech parameter mapping table in memory, mapping user selections to numerical parameters that can be directly used by generative artificial intelligence models for speech generation. For example, "gentle female voice" can be mapped to: a reference pitch lowered by a few semitones, moderate volume, and a frequency band energy distribution biased towards the low and mid frequencies; "moderate speaking speed" can be mapped to a pronunciation rate parameter of 1.0. "Mandarin Chinese" is mapped to the language code "zh-CN".

[0469] In one implementation, the server prepares prompts for voice generation through a prompt generation module. For example, when a greeting voice needs to be generated, the server can generate the following prompt: "Please use a gentle female voice to say the following in Mandarin Chinese: 'You've worked hard today, have a good rest.'" In another implementation, the server can provide text and parameters separately to the speech generation model, where prompts are primarily used to control tone and style. The generative AI model for speech generation can employ a sequence-to-sequence neural network, such as an encoder-decoder architecture. The encoder converts the text sequence into a high-dimensional representation of the speech feature sequence, and the decoder generates acoustic features (e.g., Mel spectrum) based on this representation and the speech parameter vector. These acoustic features are then converted into waveform signals via a neural vocoder.

[0470] Through the aforementioned voice parameter mapping and structured prompt statement control, the server can uniformly manage the voice expression styles in different application scenarios without modifying the business logic, thereby decoupling the voice generation module from the upper-layer application logic. At the same time, parameterized control reduces the storage and maintenance overhead of multiple fixed voice templates and improves the efficiency of computing resource utilization.

[0471] IV. Natural Language Processing and Sentiment Analysis of Textual Information and Visual Symbols In one implementation, the server uses a natural language processing module and a sentiment analysis module to jointly parse text information and visual symbols received from the terminal. The user enters text through an input box on the terminal and simultaneously selects an emoji. For example, the user enters the following on the terminal: "I've been under a lot of work pressure lately." " The terminal sends the text along with emojis to the server. The server uses its natural language processing module to perform word segmentation, part-of-speech tagging, and dependency parsing on the text, identifying words like "work," "stress," and "very" as emotion-related keywords. The server then parses the emojis into predefined sentiment tags, such as "..." This is mapped to "pain / stress".

[0472] The server incorporates multi-dimensional sentiment vector representations in its sentiment analysis module. Each input text is represented as numerical components across multiple sentiment dimensions (such as joy, sadness, anger, fear, stress, fatigue, etc.), and the server predicts these components using a sentiment analysis network. This sentiment analysis network can employ a multi-layer fully connected network or a bidirectional recurrent neural network structure. Its input includes text embedding vectors output from the natural language processing module and sentiment indicator vectors encoded from emojis. During model training, the server uses a dataset with sentiment labels, enabling the network to learn the mapping relationship from input text and emojis to sentiment vectors by minimizing cross-entropy loss or mean squared error loss. Backpropagation and gradient descent are used to update the network weights during training.

[0473] Through the aforementioned multimodal input and multidimensional sentiment representation, the server can simultaneously identify multiple emotions present in the input. For example, when a user inputs "I'm happy today, but a little tired,"... When the system detects high values ​​for both "pleasure" and "fatigue," the server can identify these emotions. Compared to traditional systems that only output a single emotion label, the emotion analysis module of this invention can fully represent complex emotions at the computational level, providing more refined feature inputs for subsequent response generation and content recommendation, thereby improving the system's accuracy in emotion recognition.

[0474] V. Generation of Emotional State-Based Prompt Statements and Response Information In one implementation, the server uses multiple emotion processing modules to generate prompts that drive the generative artificial intelligence model based on the emotion state output by the emotion analysis module and dialogue history information. The server maintains a response element library internally, which associates different emotion categories with corresponding response fragment patterns. For example, "high stress" corresponds to a combination of "comfort + actionable suggestions", and "pleasure + fatigue" corresponds to a combination of "positive experience + reminder to rest".

[0475] When the server detects that a user's emotional state is "high stress", the server can generate the following prompt: "You are a gentle and understanding psychological support assistant. The user said: 'I've been under a lot of work pressure lately.'" The user is currently feeling stressed and exhausted. Please write one or two sentences in Simplified Chinese, using a sincere and gentle tone, to comfort the user and offer practical suggestions (such as taking a break, breaking down tasks, etc.). Avoid using overly exaggerated words of encouragement. When the server detects that a user's emotional state is "happy but a little tired", the server can generate the following prompt: The user said: "I'm happy today, but a little tired." The user feels both pleased and tired. Please generate one or two sentences in Simplified Chinese that affirm their positive experience today while gently reminding them to take a break. When the server processes the aforementioned prompts using a generative AI model, the model internally encodes and decodes the prompts based on a multi-layered attention mechanism and a large-scale parameter matrix, outputting response information with high semantic richness. Before response generation, the server maps emotional states to structured prompts through a multi-emotion processing module. This process explicitly encodes emotional features at the model input, reducing the model's burden of secondary inference of implicit emotions, thereby improving the stability and consistency of response generation computationally.

[0476] VI. Multimodal Output and Content Recommendation In one implementation, the server combines the text response output by the generative artificial intelligence model with the synthesized speech output by the speech generation module through a multimodal output module to generate a composite presentation. The server then sends this composite presentation as structured data, including text fields, audio references, and optional image references, to the terminal. The terminal displays the text on its display and simultaneously plays the corresponding synthesized speech through its audio output. When a user uses a head-mounted display device, the terminal can present images of virtual entities within a virtual scene and simultaneously play the corresponding speech, thus achieving an integrated visual and auditory experience.

[0477] In terms of content recommendation, in one implementation, the server uses a content recommendation module to select content from external information providing devices or information recording devices based on emotional states. For an emotional state of "high stress," the server can select relaxation-related audio, video, or interactive content from the content database. Subsequently, the server uses a prompt generation module to construct recommendation reason prompts for the generative artificial intelligence model, such as: "Users are under a lot of work pressure and need to relax. Here is the title of a relaxing song: 'Lo-fi ChillBeats'. Please write a short paragraph of no more than two sentences in Simplified Chinese to recommend this song to the user and briefly explain why it is suitable to listen to when under pressure." The recommendation text output by the generative AI model, along with specific content links, is sent to the terminal. In this architecture, the server deeply couples sentiment state with the content selection process at the computational level, upgrading the recommendation logic from simple statistical filtering to a joint computation of "sentiment features + content features," thereby improving the personalization and interpretability of the recommendations.

[0478] VII. Technical Effects and Improvements in Computer Technology In this embodiment of the invention, the server, through attribute abstraction, multi-dimensional vectorization of emotions, and structured generation of prompt statements, maps the originally scattered business logic into a unified data processing and model invocation flow. Since all generation tasks interface with the generative AI model through prompt statements, the server can achieve rapid expansion across multiple scenarios without altering the model's internal structure by adjusting the composition and abstraction rules of the prompt statements. This reduces the maintenance costs associated with frequent template and interface modifications in traditional systems.

[0479] The server employs multimodal input and multidimensional sentiment vectors in its sentiment parsing section, enabling each input record to carry richer sentiment features and thus improving sentiment recognition accuracy. A multi-sensory processing module combines and maps various sentiments and explicitly encodes them in the prompts. This process helps generative AI models operate in a more explicit semantic space, reducing the randomness and error of the generated results. Because sentiment states are stored as vectors and labels, the server can reuse this structured data in subsequent dialogues, accelerating the response generation process and improving overall processing speed.

[0480] The server reduces the complexity of data management and model retrieval by uniformly abstracting and mapping attribute information, voice parameters, and emotional states, and storing them in memory with a unified structure. In terms of network communication, the server and terminal interact primarily based on abstracted attributes and parameters. Compared to directly transmitting large amounts of templates or pre-generated data, this reduces communication load and improves system response efficiency.

[0481] Furthermore, during the model training phase, the server jointly trains the sentiment analysis network and the auxiliary classification network. Through multi-task learning, it simultaneously optimizes sentiment recognition and related feature extraction within the same neural network, allowing the model to share feature layers during inference, improving inference efficiency and reducing computational redundancy. At the generative AI model invocation level, the server can prune and reuse prompts based on sentiment states, thereby reducing unnecessary contextual information transmission and further optimizing overall computational and communication performance.

[0482] VIII. Other Implementation Forms In another implementation, the server can extend the rules of the prompt generation module, allowing it to dynamically adjust the prompt structure based on long-term user behavior data. For example, the server can add personalized prefixes or restrictions to the prompts to constrain the output style of the generative AI model, enabling different users to receive differentiated tones and expressions.

[0483] In another implementation, the server can introduce a lightweight local model to perform some pre-parsing tasks on the terminal side, such as performing preliminary mapping of emojis or simplifying preprocessing of text, and then send the results to the server to further reduce server load and network latency.

[0484] In another implementation, the server can uniformly schedule the image generation module and the speech generation module, batch processing or prioritizing the generation tasks of multiple users to improve hardware resource utilization. The server can also cache the generation results, directly reusing existing generated content when abstract attributes and prompts are duplicated, thereby improving overall processing speed.

[0485] Through the above implementation forms, a highly efficient data flow from user input to multimodal output is formed between the server, terminal and generative artificial intelligence model. At the technical level, the system realizes unified management of attribute abstraction, sentiment analysis, multi-sentiment synthesis and content recommendation. It not only supports human dialogue and soothing behavior, but also improves the traditional computing system in terms of computing architecture, data structure and model calling strategy, thereby achieving significant technical effects in terms of processing speed, recognition accuracy, communication load and resource utilization efficiency.

[0486] use Figure 14 The processing procedure is explained.

[0487] Step 1: Users input or select attribute information and voice output options for virtual entities through a graphical interface on the terminal. Input includes: appearance attributes (e.g., hair color, eye color, body type, clothing style, etc.) and voice preferences (e.g., language, timbre, speech rate, gender, etc.). The terminal encapsulates this information into structured data (e.g., a set of key-value pairs) and sends it to the server via the communication module. Output consists of attribute information data and voice selection data sent to the server.

[0488] Step 2: The server receives attribute information data and voice selection data from the terminal. The input is structured data sent by the terminal. The server parses this data in the data management module, converting JSON or other data formats into an internal object structure, and performs field validation and completion (e.g., filling in default values). The server outputs internal attribute information objects and voice selection objects for subsequent processing.

[0489] Step 3: The server performs attribute abstraction processing based on internal attribute information objects. The input is specific attribute information (e.g., "blue hair," "green eyes," "athletic body type," etc.). The server queries a pre-built attribute ontology table or hierarchical label mapping table to map specific attributes to higher-level conceptual attributes (e.g., "cool-toned hair color," "bright-colored eyes," "low-fat body type," etc.), transforming the original attribute vector into an abstract attribute vector in terms of data structure. Through this data processing, the server reduces redundant details and outputs a set of abstract attribute information.

[0490] Step 4: The server generates image-based prompts to drive generative AI models based on abstract attribute information. The input is a set of abstract attribute information. In the prompt generation module, the server calls string template combination logic to concatenate several higher-level conceptual attributes into a natural language description in a predetermined order. For example, it generates the prompt: "Blonde, blue-eyed, adventurer-style anime / manga character, full-body image, high-quality illustration style." The server outputs a text-based prompt for image generation.

[0491] Step 5: The server inputs image-based prompts into the generative AI model interface for image generation. The input consists of the prompt text and optional generation parameters (resolution, number of steps, etc.). The server encodes the prompts into the format required by the model interface and sends them to the image generation model (e.g., a diffusion model). Internally, the image generation model performs multi-step noise inversion and feature calculations, and the server receives the output pixel matrix or image file data. The server performs format conversion and compression on the image data (e.g., from floating-point matrix to PNG or JPEG encoding), and the output is a visual representation of a virtual entity that can be displayed on the terminal.

[0492] Step 6: The server maps speech parameters to the selected information related to the speech output. The input is the user's selected speech preferences (language, gender, speech rate, timbre, etc.). The server queries the speech parameter mapping table, converts each selection into numerical or categorical parameters (such as pitch offset, speaking rate, volume level, language encoding, etc.), and combines them into a complete speech parameter vector. During data processing, the server performs range pruning and default value filling, and the output is a standardized set of speech parameters.

[0493] Step 7: After entering the chat interface on the terminal, the user inputs text information and selects visual symbols (such as emoticons). The input is natural language text (e.g., "I've been under a lot of work pressure lately"). The terminal encodes the text and emojis into a unified string or structured message and sends it to the server via the communication module. The output is the chat message data sent to the server.

[0494] Step 8: The server uses a natural language processing module and a sentiment analysis module to parse chat message data. The input is message text containing text and emoticons. The server first calls word segmentation, part-of-speech tagging, and dependency parsing algorithms to break the text down into words and label their syntactic relationships; then it parses the emoticons and maps them to predefined sentiment tags. The server combines text and emoticon features into a feature vector, which is then input into the sentiment analysis network to perform forward inference, obtaining multi-dimensional sentiment vectors (e.g., numerical values ​​for joy, sadness, stress, etc.) and the main sentiment category. The server outputs numerical and categorical sentiment state data.

[0495] Step 9: The server generates response strategy information based on sentiment state data and dialogue history. Inputs include the current sentiment state vector, the primary sentiment category, and historical dialogue records. The server searches for response element templates (e.g., "comfort + specific advice," "affirmation + reminder to rest," etc.) corresponding to the primary and secondary sentiments in the multi-sensory processing module, and determines which response content to retain or avoid based on the historical context. The server determines the set of response elements required for this round of responses through a rule engine or weighted combination algorithm, and outputs a structured response strategy (including element type, sentiment weight, etc.).

[0496] Step 10: The server generates prompts for text generation based on response strategies and emotional states. Inputs include a set of response elements, emotional tags, and the user's original input. The server embeds this information into natural language instructions through the prompt generation module, forming detailed requirements for the generative AI model. For example, when the emotional state is "high stress," the server generates the prompt: "You are a gentle, understanding psychological support assistant. The user says: 'I've been under a lot of work stress lately.'" The user is currently feeling stressed and exhausted. Please write one or two sentences in Simplified Chinese, using a sincere and gentle tone, to comfort the user and offer practical suggestions (such as taking a break, breaking down tasks, etc.). Avoid using overly exaggerated words of encouragement. The server output will be the prompt statement used for text generation.

[0497] Step 11: The server sends the aforementioned text-based prompt to the generative AI model used for text generation and retrieves the response text. The input is the prompt text. The server encodes the prompt via the model interface module and sends an inference request to the generative AI model. The generative model internally performs multi-layer attention calculations and sequence prediction, and the server receives the natural language response text output by the model. The server performs post-processing as necessary, such as removing redundant prefixes and suffixes and standardizing punctuation, and outputs the user-oriented response text information.

[0498] Step 12: The server generates prompts or configurations for speech generation based on a set of voice parameters and response text. The inputs are standardized voice parameters and response text. The server selects an appropriate speech generation mode: in one mode, the server generates prompts in natural language, such as "Please speak the following in a gentle female voice in Mandarin Chinese: 'You've worked hard today, take a good rest.'"; in another mode, the server uses the response text and voice parameters in a structured manner as input configuration for the TTS model. The server output is speech generation request data, including text content and voice parameters.

[0499] Step 13: The server invokes a generative artificial intelligence model for speech generation to produce synthesized speech data. The input is the speech generation request data. The server sends a request to the speech synthesis model via the speech generation interface module. Internally, the model uses an encoder-decoder structure to convert the text into acoustic features and generates waveform data using a vocoder. The server encodes and compresses the waveform (e.g., converting it to MP3 or other formats) and standardizes parameters such as volume and level. The server output is playable audio data.

[0500] Step 14: The server combines image data (visual representation of virtual entities), text response information, and synthesized speech data in its multimodal output module. Inputs include image reference information, text responses, audio data, and recommended content. The server constructs a composite response object containing multiple fields, including a text string for display, audio binary or link address for playback, and image URL or image data for rendering. The server outputs the composite response object and sends it to the terminal via the communication module.

[0501] Step 15: The terminal receives a composite response object from the server. The input is structured data containing text, audio, and image information. The terminal parses this data, passing the text response to the display module for presentation as a speech bubble on the screen; passing the audio data to the audio playback module for playback via speakers or headphones; and, when necessary, loading a virtual entity image into the character area of ​​the interface for display. The terminal output is a multimodal presentation effect on the user interface.

[0502] Step 16: Users observe virtual entity images, read text responses, and listen to synthesized speech on the terminal. After subjectively evaluating the current feedback, they can continue to input new text or emoticons. The input consists of new user interaction content. The terminal then sends the new input to the server, thus forming a cyclical dialogue and generation process. The output is new chat message data, which enters subsequent similar processing flows.

[0503] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0504] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0505] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.

[0506] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0507] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.

[0508] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.

[0509] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.

[0510] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0511] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.

[0512] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0513] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0514] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0515] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0516] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0517] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0518] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.

[0519] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0520] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0521] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0522] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0523] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0524] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0525] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0526] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.

[0527] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0528] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.

[0529] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.

[0530] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.

[0531] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0532] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.

[0533] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0534] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0535] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0536] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0537] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0538] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0539] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.

[0540] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".

[0541] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0542] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0543] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0544] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0545] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0546] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0547] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.

[0548] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12 to analyze the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12 to generate a menu using a generation AI. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12 to provide the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0549] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.

[0550] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.

[0551] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.

[0552] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0553] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.

[0554] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0555] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by a perspective equivalent to the field of vision of an average healthy person).

[0556] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0557] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0558] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0559] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0560] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0561] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.

[0562] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".

[0563] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0564] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0565] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0566] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0567] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0568] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0569] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.

[0570] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0571] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.

[0572] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see [reference]). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.

[0573] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.

[0574] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.

[0575] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).

[0576] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.

[0577] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."

[0578] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values ​​representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.

[0579] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).

[0580] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.

[0581] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0582] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.

[0583] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.

[0584] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.

[0585] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.

[0586] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.

[0587] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.

[0588] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.

[0589] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.

[0590] In addition, the following notes are provided in response to the above explanation.

[0591] Example 1 (Note 1) An information processing system, characterized in that it comprises: An apparatus for receiving input information containing attribute information from a user terminal, regularizing the attribute information into feature information of a higher-level concept, and generating a prompt statement for instructing the generation of a visual representation based on the feature information. An apparatus for converting the prompt statement into a form suitable for input into a generative artificial intelligence model, determining the type of the generative artificial intelligence model, and selecting a generative artificial intelligence model for image generation or a generative artificial intelligence model for text generation based on the determination result. A device for inputting the converted prompt statement and image generation parameters into a generative artificial intelligence model for selected image generation to generate visual representation data reflecting the feature information, and converting the visual representation data into image data that can be sent to a user terminal and sending it. A device for determining whether an audio representation generation request exists based on the user's selection information; if the determination result is positive, generating a prompt statement and audio generation parameters for input into the generative artificial intelligence model for audio generation; generating audio data through the generative artificial intelligence model for audio generation; and sending the audio data to the user terminal. An apparatus for obtaining text and symbol information from user terminal text and symbol information received based on a session, analyzing the text and symbol information using natural language processing and sentiment analysis techniques to classify the user's emotional state, determining a support strategy based on the emotional state, generating a prompt statement containing reassuring response content based on the support strategy, and inputting the prompt statement into a generative artificial intelligence model for text generation to generate a response message. A device for generating prompt statements for controlling a generative artificial intelligence model for audio generation to generate audio data corresponding to the response message, generating the audio data when necessary, and prompting the response message and the audio data to a user terminal; A history management device for associating and recording the feature information, the prompt statements, the visual representation data, the audio data, and the emotional state, and for providing users with reusable prompt statement candidates or visual representation candidates based on the records.

[0592] (Note 2) The information processing system according to Appendix 1 is characterized in that, The visual representation data includes a virtual object image that integrates multiple feature information specified by the user, and generates a correction prompt statement using past prompt statements or feature information obtained by the history management device, and presents the correction prompt statement to the user terminal so that the user can make adjustments based on the existing virtual object image.

[0593] (Note 3) The information processing system according to Appendix 1 is characterized in that, During the generation of the audio data, audio generation parameters, including timbre type, speech rate type, and performance style type, are automatically selected based on the user's emotional state and the feature information. The audio generation parameters are then used to control the generative artificial intelligence model for audio generation, so as to generate an audio representation consistent with the visual representation data and sentiment analysis results.

[0594] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for acquiring attribute information input by a user through an information terminal, receiving the attribute information through a communication method, performing higher-level concept normalization processing on the attribute information, and generating prompt statements to instruct generative artificial intelligence models to generate visual information of virtual representations. An apparatus for generating voice generation control parameters or voice prompt statements based on voice attribute information selected by a user through an information terminal and the attribute information, and for generating prompt statements to instruct a generative artificial intelligence model to generate voice information. A device for performing natural language processing parsing on text and emoji information received from a messaging terminal in the form of chat, and extracting emotional information for typifying and representing the user's emotional state. An apparatus for constructing message candidates for providing a sense of reassurance based on the emotional information and the attribute information, and for generating prompt statements for instructing a generative artificial intelligence model to generate at least a portion of the message candidates as natural language messages. An apparatus for providing in real time the visual information of the virtual representation generated by the generative artificial intelligence model and the natural language message, as well as the speech information generated by the generative artificial intelligence model for speech generation, to an information terminal so that the information terminal can display a virtual representation reflecting the individuality of the user in a virtual space and play the speech. An apparatus for sequentially acquiring changes made by the user to the attribute information and / or the voice attribute information, and providing the visual and voice information of the virtual representation in a continuously updated state over time by re-executing the prompt statement generation process and the generative artificial intelligence model generation process.

[0595] (Note 2) The information processing system according to Appendix 1 is characterized in that, The visual information of the virtual representation is image data generated by inputting prompt statements based on the attribute information into a generative artificial intelligence model for image generation, and the prompt statements are constructed as a superordinate concept description including body information, appearance color information, expression information, and clothing information.

[0596] (Note 3) The information processing system according to Appendix 1 is characterized in that, The voice information is generated by inputting voice prompts or voice control parameters that reflect the user's emotional information and voice quality attributes into a generative artificial intelligence model for voice generation. The voice quality attributes are constructed as a higher-level conceptual description including voice type information, age perception information, speech rate information, and emotional expression information.

[0597] Example 2 (Note 1) An information processing system, characterized in that it comprises: An apparatus for acquiring attribute information and voice parameters related to speech generation from a user terminal, receiving the attribute information and voice parameters as input, abstracting the attribute information and voice parameters into higher-level feature information, and generating prompt statements to instruct a generative artificial intelligence model to generate virtual objects based on the higher-level feature information; An apparatus for parsing the speech parameters into control information representing pitch, speech rate, speech manner, speech style and emotional state, and using the control information to generate prompt statements for instructing a generative artificial intelligence model to generate speech information with predetermined speech content and predetermined speech features. A device for performing natural language processing on textual and symbolic information acquired in the form of a conversation to parse high-level emotional state information, and generating prompt statements based on the emotional state information to instruct a generative artificial intelligence model to generate response information that gives the prompted subject a sense of reassurance, and a device for prompting the prompted terminal with the response information obtained from the generative artificial intelligence model. And an apparatus for sending the voice parameters and the prompt statement to a voice generation computing device via a communication path, and distributing the voice information obtained from the voice generation computing device to the prompted terminal for regeneration.

[0598] (Note 2) The information processing system according to Appendix 1 is characterized in that, The device for generating prompt statements is further configured to: generate prompt statements based on the attribute information and the higher-level feature information, instructing the image generation to use a generative artificial intelligence model to generate image information with virtual object appearance features, and control the display of the image information to the prompted terminal.

[0599] (Note 3) The information processing system according to Appendix 1 is characterized in that, The apparatus for generating prompt statements and generating voice information is further configured to: input the prompt statements generated from the voice parameters and the spoken content into a generative artificial intelligence model for voice generation to generate the voice information; obtain evaluation information of the prompted subject regarding the regeneration result of the voice information; update the voice parameters to high-level control information based on the evaluation information; and regenerate the prompt statements based on the updated control information, thereby generating voice information that conforms to the intent of the prompted subject in an iterative manner.

[0600] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: An apparatus for acquiring user attribute information from a user terminal, converting the attribute information into abstract attribute information based on a higher-level concept, and generating prompt statements based on the abstract attribute information to enable a generative artificial intelligence model to generate a visual representation of a virtual entity. A device for acquiring selection information related to voice output from a user terminal, mapping the selection information to voice parameters of a higher-level concept, and generating prompt statements for enabling the voice generation to use a generative artificial intelligence model to generate synthesized speech and instructing the generation of the synthesized speech based on the voice parameters and text information. A device for receiving text information and visual symbols from a user terminal in a chat format, converting the text information and visual symbols into higher-level conceptual language information and emotional indicators, using natural language processing technology and sentiment analysis technology to quantify and classify the user's emotional state, and generating prompt statements for enabling a generative artificial intelligence model to generate response information and support information to provide the user with a sense of reassurance based on the quantified and classified emotional state and dialogue history information. A device for converting the response information and support information output from the generative artificial intelligence model into a composite presentation form of a higher-level concept including text information and voice information, and sending it to the user terminal for presentation; An apparatus for acquiring content of a higher-level concept from an external information providing device or information recording device based on the emotional state, selecting the content, generating prompt statements for enabling a generative artificial intelligence model to generate recommendation reasons for the content, and presenting the content and the recommendation reasons to the user terminal; An apparatus for simultaneously extracting multiple emotions using natural language processing techniques when the emotional state contains multiple emotions, and generating prompt statements for enabling a generative artificial intelligence model to generate a comprehensive response message consisting of response elements corresponding to each emotion.

[0601] (Note 2) According to the information processing system described in Appendix 1, the device for obtaining user attribute information from a user terminal and generating a prompt statement for enabling a generative artificial intelligence model to generate a visual representation of a virtual entity is configured to generate a prompt statement based on the user attribute information for enabling a generative artificial intelligence model composed of an image generation device of a higher-level concept to generate a visual representation of a virtual entity, and to generate a visual representation of the virtual entity using the generative artificial intelligence model.

[0602] (Note 3) According to the information processing system described in Appendix 1, the device for generating a prompt statement based on speech parameters and text information to enable the speech generation to generate synthesized speech using a generative artificial intelligence model and instructing the generation of the synthesized speech is configured to generate a prompt statement based on user speech output-related selection information to enable the generative artificial intelligence model, which is composed of a speech synthesis device of a higher-level concept, to generate a speech representation, and to generate the speech representation using the generative artificial intelligence model.

Claims

1. An information processing system, characterized in that, include: processor; The processor is configured to: receive user-specified attribute information as input, generate prompt information to instruct the artificial intelligence model to generate a virtual character; and generate prompt information to instruct the artificial intelligence to generate speech based on the user's selection. The system parses the text and emoticons received in the chat, analyzes the user's emotions using natural language processing technology, and generates prompts based on the analysis results to instruct the AI ​​model to generate messages that provide a sense of reassurance. The generated messages are then presented to the user.

2. The information processing system according to claim 1, characterized in that, The processor is configured to generate prompts based on user attribute information to indicate the generation of a virtual character, and to generate the virtual character using image generation artificial intelligence.

3. The information processing system according to claim 1, characterized in that, The processor is configured to generate prompts for instructing the generation of speech based on the user's selection, and to generate the speech using speech generation artificial intelligence.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A