Information processing system

CN122797484APending Publication Date: 2026-09-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610319385.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-19
Filing Date
2026-03-16
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

当用户希望设计一个全新或复杂的角色人格时,由于缺乏参考信息或典型行为模式范例,往往需要大量试错和反复修改,效率低且创作门槛较高

Benefits of technology

服务器在本发明中不仅执行传统的“接收文本—返回文本”的简单流程,而是通过以下机制实现对计算机技术本身的改进:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122797484A_ABST
    Figure CN122797484A_ABST
Patent Text Reader

Abstract

The application provides an information processing system, comprising a processor configured to: parse information from a user by using a generative artificial intelligence model, and automatically generate a personality prompt for indicating a specific behavior or reaction; make the generated personality prompt editable or appendable by the user through an interface, so as to customize the personality prompt; and monitor an emotional state of the user in real time, and dynamically adjust the personality prompt according to the emotional state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology disclosed herein relates to an information processing system. Background Technology

[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot response to the user's speech.

[0003] With the development of generative artificial intelligence technology, the demand for using generative AI models (such as large-scale language models) to design personality settings for virtual characters, dialogue agents, or interactive systems is constantly increasing. However, several prominent problems in existing technologies remain unresolved.

[0004] First, in existing systems, when generating personality-related prompts using generative AI models, the process often relies solely on static user input or pre-defined role descriptions for a one-time generation, lacking sufficient consideration of individual user differences and emotional states. This results in personality prompts that fail to consistently align with the user's current psychological state and interaction needs, thus impacting the user experience and sense of role immersion.

[0005] Secondly, existing character personality settings or personality prompts are mostly written manually by developers or professional creators. Ordinary users, without experience, find it difficult to efficiently conceive clearly structured and behaviorally defined personality prompts. Even when some systems introduce automatic generation functions, they often lack convenient editing and expansion mechanisms, preventing users from easily personalizing the automatically generated content and limiting their creative freedom.

[0006] Furthermore, existing technologies typically do not provide systematic assistance to address the "lack of inspiration" problem users face when brainstorming new character ideas. When users wish to design a completely new or complex character, the lack of reference information or typical behavioral patterns often necessitates extensive trial and error and repeated revisions, resulting in low efficiency and a high creative threshold.

[0007] In summary, the problem this invention aims to solve is to provide a system that can automatically generate personality prompts using a generative artificial intelligence model, support users in flexibly customizing personality prompts through an interface, dynamically adjust personality prompts based on the user's emotional state, and provide reference information and creative inspiration when the user is conceiving new personality prompts, thereby improving the relevance, personalization, and creative efficiency of personality prompts and enhancing the overall user interaction experience. Summary of the Invention

[0008] To address the aforementioned issues, this invention provides an information processing system that, by introducing a generative artificial intelligence model, a user interface, and an emotion state monitoring and dynamic adjustment mechanism, enables the automatic generation, customizable editing, and adaptive adjustment of personality prompts, and provides reference information and inspiration support when users are conceiving new personality prompts.

[0009] Specifically, according to one aspect of the present invention, an information processing system is provided, characterized in that it includes a processor; the processor is configured to: analyze information from a user using a generative artificial intelligence model, and automatically generate personality prompts to indicate specific behaviors or reactions. Through this configuration, the system can automatically abstract and output multiple personality prompts describing typical behavioral patterns and reaction methods of a character based on user-inputted information such as role name, personality traits, and scene settings, thereby reducing the workload of users manually writing personality profiles.

[0010] According to another aspect of the invention, the processor is further configured to allow the generated personality prompts to be edited or appended by the user through an interface, thereby customizing the personality prompts. By presenting the automatically generated personality prompts in an editable form within the human-computer interaction interface, users can modify, add, delete, and reorganize the text content to form a customized set of personality prompts that conforms to their personal creative intentions and specific application scenarios, thus ensuring generation efficiency while taking into account the personalization and subtle expression of personality settings.

[0011] According to another aspect of the invention, the processor is further configured to: monitor the user's emotional state in real time and dynamically adjust the personality cues based on the emotional state. To this end, the processor can infer the user's emotional state based on the user's current interaction content, tone of voice, word choice, or other available emotion recognition signals, and when generating or invoking personality cues, adjust the tone, style, or behavioral tendencies of the cues according to the detected emotional state, making the character's behavior and reactions more closely resemble the user's current psychological state, thereby enhancing the empathy and immersion of the interaction.

[0012] Furthermore, according to embodiments of the present invention, the processor can be configured to generate prompts using a generative artificial intelligence model to instruct the behavior or reactions of a specific character. By incorporating character characteristic information such as character identity, backstory, and occupation type when generating prompts, the processor can output personality prompts that are highly consistent with the character's settings, making the character's behavior in dialogues or stories more consistent and predictable.

[0013] Furthermore, according to an embodiment of the present invention, the processor can also be configured to: provide reference information to the user when the user conceives new personality prompts using a generative artificial intelligence model. The reference information may include typical character archetypes, common behavioral patterns, emotional response templates, dialogue style examples, etc. When the user requests creative inspiration on the interface, the processor invokes the generative artificial intelligence model to output several reference prompts or example sentences, which the user can rewrite, combine, or extend based on these references, thereby effectively reducing the difficulty of conception while ensuring creative freedom.

[0014] Through the aforementioned technical means, the system provided by this invention can achieve the following effects: Firstly, it significantly reduces the initial construction cost of personality prompts by utilizing generative artificial intelligence models; secondly, it ensures a high degree of customizability of personality prompts by leveraging the editing and appending functions provided by the user interface; thirdly, it makes the personality prompts more psychologically aligned with the user's emotional state through real-time monitoring and dynamic adjustment; and fourthly, it effectively assists users in conceiving new personality prompts by providing reference information when they lack inspiration. Therefore, the system of this invention can effectively solve the problems of cumbersome personality prompt generation and customization processes, lack of emotional adaptation, and creative support in existing technologies.

[0015] "System" refers to a hardware and software complex consisting of at least one processor and associated storage devices, communication interfaces, and user interfaces, used to perform personality prompt generation, editing, adjustment, and related control processing.

[0016] A processor is a computing unit that can execute program instructions, perform calculations, logical judgments, and flow control on input data to achieve functions such as personality prompt generation, editing control, emotion monitoring, and data storage management. It includes, but is not limited to, a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or any combination thereof.

[0017] "Generative AI models" refer to AI models trained on massive amounts of data that can automatically generate text and other content based on input information, including but not limited to large-scale language models, dialogue models, or other generative models that can output natural language prompts.

[0018] "User" refers to an entity that provides input information to the system, receives system output, and views or edits personality prompts through a terminal or other interactive device. It can be a natural person or a virtual entity authorized in the system.

[0019] "Information" refers to various data content provided by users to the system and parsed by generative artificial intelligence models, including but not limited to natural language text input, option selection, parameter settings, historical interaction records, and other valid data used to describe the needs of a role or individual.

[0020] "Personality cues" refer to a set of textual instructions or descriptive statements that are automatically generated by the system or edited by the user to instruct a specific role or system on the behavior, reaction, tone, attitude, or decision-making style that should be exhibited in a given scenario.

[0021] "Interface" refers to a visual or non-visual interactive layer used for information exchange between users and systems, including but not limited to graphical user interfaces (GUI), web interfaces, mobile application interfaces, command-line interfaces, and interactive interfaces that achieve input and output through voice, gestures, etc.

[0022] "Editing" refers to the user's modification of the generated personality prompts through the interface, including but not limited to adding, deleting, replacing, rewriting, splitting, merging, and adjusting the prompt structure or expression.

[0023] "Adding" refers to the user's action of adding one or more personality cues to the existing set of personality cues in order to expand the scope of the description of the character's behavior or reaction.

[0024] "Customization" refers to the user's ability to edit, add, delete, rearrange, or tag automatically generated personality prompts, thereby creating a set of personalized personality prompts that meet the user's individual needs and specific application scenarios.

[0025] "Emotional state" refers to the psychological and emotional tendencies expressed or implied by a user at a specific point in time, including but not limited to pleasure, sadness, anger, tension, calmness, excitement, etc., which can be inferred or identified through text content, voice features, physiological signals or other available clues.

[0026] "Real-time monitoring" refers to the system continuously or quasi-continuously analyzing user-related data at predetermined time intervals or when triggering conditions occur during user interaction, in order to update the judgment of the user's emotional state in a relatively short period of time.

[0027] "Dynamic adjustment" refers to the system automatically modifying or regenerating parameters such as the content, tone, level of detail, or behavioral tendencies of personality prompts when it detects changes in the user's emotional state or other related conditions, so that the output results adapt to the user's current needs as time and context change.

[0028] "Specific behavior or response" refers to the specific manifestations of actions, language expressions, attitude choices, or decision results that a role or system is expected to exhibit under given scenarios, situations, or input conditions.

[0029] "Specific role" refers to a virtual character, dialogue agent, game character, or other subject that can be described in a personified way and has relatively clear settings. The role may include attribute information such as name, identity, background, personality traits, and values.

[0030] "Reference information" refers to the content provided by the system to inspire or assist in the creation of new personality prompts when users are conceiving or editing them. This includes, but is not limited to, descriptions of typical character archetypes, examples of behavioral patterns, examples of dialogue styles, templates for emotional responses, and suggestive text related to personality settings. Attached Figure Description

[0031] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.

[0032] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.

[0033] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.

[0034] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.

[0035] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.

[0036] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and head-mounted terminal according to the third embodiment.

[0037] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.

[0038] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.

[0039] Figure 9 This represents an emotion map that maps multiple emotions.

[0040] Figure 10 This represents an emotion map that maps multiple emotions.

[0041] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.

[0042] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.

[0043] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.

[0044] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation

[0045] Hereinafter, an example of an implementation of the system according to the present disclosure will be described with reference to the accompanying drawings.

[0046] First, let me explain the terminology used in the following instructions.

[0047] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.

[0048] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.

[0049] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.

[0050] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.

[0051] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.

[0052] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.

[0053] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.

[0054] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0055] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.

[0056] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.

[0057] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0058] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.

[0059] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.

[0060] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0061] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0062] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.

[0063] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.

[0064] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."

[0065] Existing personality prompt generation technologies based on generative artificial intelligence models typically simply concatenate the user-inputted role description into a simple prompt text, which is then output by the generative AI model. This approach suffers from the following technical problems: First, the server lacks systematic and structured preprocessing of the role information from the terminal, failing to accurately extract key elements such as character attributes, behavioral tendencies, and activity environments. This results in redundant and ambiguous data representation input to the model, leading to insufficient controllability in the behavioral rules and response methods of the generated personality prompts, thus affecting the consistency and stability of the generated results. Second, when constructing model input, the server usually only constructs a single-level natural language instruction, failing to distinguish between system-level and user-level prompts. This makes it difficult to finely define the behavioral boundaries and contextual focus of the generative AI model, easily leading to problems such as role setting drift and unclear values. Third, the server lacks a computer-oriented post-processing mechanism for the prompts output by the generative AI model, lacking automated content detection, invalid information deletion, and format normalization processes. This results in noise in the generated results, hindering their automatic invocation and reuse by other applications. Fourth, servers generally do not dynamically learn and generate conditions based on users' operation history and candidate selection results, nor do they have the ability to adjust the content and detail of prompt statements in real time based on users' emotional state information. As a result, they cannot adaptively optimize the generation process of personality prompt statements for different users and in different situations at the system level.

[0066] From a computer technology perspective, the aforementioned problems directly lead to fragmented processing flows on the server side in key stages such as data preprocessing, model input construction, model output post-processing, and personalized generation control. The generation pipeline lacks a closed-loop optimization mechanism, failing to fully unleash the potential of generative AI models and struggling to maintain high-quality, controllable, and consistent personality prompts in scenarios with large-scale user access. Therefore, it is necessary to propose a personality prompt generation system that performs structured preprocessing of character input information on the server side, hierarchically constructs model prompts, automatically post-processes model output, and adaptively optimizes based on user operation history and emotional state. This system aims to comprehensively improve the computer's processing efficiency, generation quality, and controllability in such tasks.

[0067] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.

[0068] In this invention, the server includes a preprocessing device for receiving input information including character attribute information and behavioral information from the terminal and performing normalization and structuring preprocessing; a generation device for generating model input text containing system-level prompts and user-level prompts based on structured data; a prompt generation device for calling a generative artificial intelligence model to automatically generate personality prompts that specify character value standards, behavioral rules, response methods, and language styles based on the model input text; a postprocessing and communication device for performing content detection, invalid information deletion, and formatting on the generated personality prompts and sending them to the terminal; an update device for receiving user editing information or regeneration instructions and regenerating or updating the prompts accordingly; a learning device for learning prompt generation conditions based on user operation history information and selection history information and reflecting them in subsequent prompt generation to output optimized personality prompts for each user; and an adjustment device for dynamically adjusting the content or level of detail of the generated personality prompts based on user emotional state-related status information. This allows for the formation of a complete technical chain on the server side, encompassing input information structuring, hierarchical prompt construction, generative AI model reasoning, post-processing of results, and personalized and context-adaptive control. This reduces the uncertainty caused by directly inputting unstructured text into the model, improves the controllability and consistency of personality prompts in terms of behavioral constraints and language style, and enables automatic optimization of the prompt generation process for different users and usage scenarios through learning and adjusting user history and emotional state. Consequently, this substantially improves the processing performance and generation quality of computers in personality prompt generation tasks.

[0069] "Terminal" refers to an information processing device operated by a user, equipped with input and output functions, and interacting with a server through a communication network, including but not limited to computer equipment, mobile communication equipment, tablet equipment, or other software carriers capable of running client programs or browsers.

[0070] A "server" refers to an electronic computing device or computer system that is equipped with information processing programs to receive terminal requests, perform data processing, generate results, and return the results to the terminal. It can consist of a single computer device or multiple computer devices.

[0071] "Information processing device" refers to a data processing functional unit implemented on a server. It can be implemented in hardware, software, or a combination of hardware and software, and is used to analyze, transform, generate, and control the output of received data.

[0072] "Input information" refers to the data content sent by the terminal and received by the server, which includes at least the character's attribute information and behavioral information, and may also include environmental information, user preference information, emotional state information, and other descriptive data related to the character's settings.

[0073] "Character attribute information" refers to data information used to describe the basic characteristics of a character, including but not limited to personality traits, social roles, occupational types, age range, worldview background, value orientation, and other abstract or specific attributes related to the character setting.

[0074] "Behavioral information" refers to data that describes a person's behavioral patterns, tendencies, or reactions in a specific situation, including but not limited to action strategies, decision-making preferences, and typical reactions to others or events.

[0075] "Preprocessing" refers to the normalization, cleaning, classification, and structuring operations performed on input information before it is provided to a generative artificial intelligence model, in order to transform unstructured or semi-structured natural language descriptions into structured data forms that are more suitable for subsequent model processing.

[0076] "Structured data" refers to a data representation that has been preprocessed and organized according to predetermined fields or attribute categories. It typically uses key-value pairs, objects, tables, or other data structures with clearly defined fields to facilitate the retrieval and combination of different elements by programs.

[0077] "Natural language instruction text" refers to text content written in natural language used to explain task objectives, constraints, background information, and output requirements to generative artificial intelligence models, including system-level prompts and user-level prompts.

[0078] "System-level prompts" refer to natural language text that serves as the overall behavioral guidelines for generative artificial intelligence models. They are used to specify the roles, styles, constraints, and global rules that the model should follow during dialogue or generation.

[0079] "User-level prompts" refer to natural language text generated from user input or directly related to user intent, based on specific tasks or roles, used to provide specific contextual explanations or output requests to generative artificial intelligence models.

[0080] "Model input text" refers to natural language text or its encoded representation, which is composed of system-level prompts and user-level prompts and is used as input to generative artificial intelligence models. It explicitly includes task descriptions, role settings, and generation requirements.

[0081] "Generative AI models" refer to AI models that are trained on large-scale datasets and can automatically generate new text from input text, including but not limited to sequence generation models based on deep neural networks and self-attention mechanisms.

[0082] "Prompt statements" refer to natural language text used to guide generative artificial intelligence models in generating target content, and include at least personality prompt statements that specify a person's value standards, behavioral rules, response methods, and language style.

[0083] "Personality cue statements" refer to cue statements that describe the consistency settings of a specific person in terms of values, behavioral principles, emotional reactions, and language style in natural language form, and are used to constrain or guide the relevant output of that person in generative artificial intelligence models.

[0084] "Post-processing" refers to the process performed by the server after the generative artificial intelligence model outputs a prompt statement. This includes content detection, deletion of invalid information, formatting, and necessary filtering and correction operations to obtain the final output text that conforms to predetermined specifications.

[0085] "Editing information" refers to the operations performed by users on the terminal when modifying, supplementing, or partially replacing prompts generated by the server, including adding, deleting, and modifying text, as well as adjusting style or constraints.

[0086] "Regenerate instruction" refers to a request sent by a user to a server through a terminal, which is used to re-invoke a generative artificial intelligence model to generate new prompts while keeping all or part of the input information unchanged.

[0087] "Generation conditions" refers to the set of parameters, constraints, and strategies used to control the generation results during the prompt statement generation process, including but not limited to character attribute weights, output length, style preferences, level of detail, and configurations related to user preferences.

[0088] "Operation history information" refers to user operation behavior data recorded during system use, including but not limited to the user's selection of candidate prompt statements, editing behavior, number of regenerations, and logs related to interface interaction.

[0089] "Selection history information" refers to the historical records formed when a user selects from multiple candidate prompts or multiple options, which are used to reflect the user's preference for different prompt styles or settings.

[0090] "Learning device" refers to a functional unit used to model, adjust and optimize the generation conditions based on operation history information, selection history information and prompt statements. It can be implemented through machine learning algorithms or rule update mechanisms.

[0091] "Emotional state" refers to the state information that reflects a user's psychological or emotional state within a specific time period, including but not limited to emotional categories or corresponding intensity indicators such as tension, pleasure, frustration, and calmness, which can be input by the user or inferred by the system.

[0092] "Status information" refers to data related to a user's emotional state that can be collected by the terminal and sent to the server, including explicit self-emotion reports from the user, evaluation information of content, and emotional tags inferred through interaction behavior.

[0093] "Adjustment device" refers to a functional unit that dynamically adjusts the content, tone, level of detail, or length of the generated personality prompts based on state information, so as to achieve adaptive support for user needs in different emotional states.

[0094] "Storage device" refers to a storage medium or storage system used to store prompt statements, character attribute information, behavior information, operation history information, and data related to statistical analysis and template generation, and may include volatile storage media and non-volatile storage media.

[0095] "Statistical information" refers to quantitative data results about character attributes, behavioral patterns, user preferences, and the usage of prompt statements obtained through statistical analysis based on data accumulated in storage devices.

[0096] "Recommended templates" refer to reusable personality prompt statement structures or text fragments extracted based on statistical information and historical generation results, used to provide reference or serve as an initial framework for subsequent prompt statement generation.

[0097] The embodiments of this invention will be described using a typical client-server architecture as an example. However, those skilled in the art will understand that various modifications can be made to the hardware configuration, software framework, and network topology without departing from the spirit of this invention. The core of this invention lies in the following: the server performs structured preprocessing on natural language character descriptions from the terminal, constructs hierarchical prompt statements as input to a generative artificial intelligence model, and then performs technical post-processing and personalized adjustments after obtaining the personality prompt statements, thereby forming an efficient, controllable, and learnable personality prompt statement generation pipeline within the computer.

[0098] A server is implemented on one or more electronic computing devices, such as a computer system running a general-purpose operating system (like a UNIX-like operating system). A server may include at least one central processing unit (CPU), an optional graphics processing unit (GPU) or dedicated acceleration chip (such as a deep learning accelerator), main memory, non-volatile memory, and a network interface card. At the software level, the server may run web application frameworks (such as Python-based web frameworks), natural language processing libraries (such as word segmentation and feature extraction libraries), and deep learning inference frameworks (such as model inference libraries supporting the Transformer architecture).

[0099] A terminal can be a user-held device with display and input capabilities, such as a smartphone, tablet, or general computer terminal. The terminal runs client applications or a browser and establishes a connection with the server via a secure communication protocol. The terminal provides a graphical user interface, enabling users to input personality attribute and behavioral information, view personality prompts returned by the server, and edit or select them.

[0100] Users describe the target character in natural language through the input interface on the terminal. For example, users can enter "brave, values ​​honor, willing to protect the weak" in the personality description field, "knight" in the role type field, and "medieval fantasy world" in the worldview field. They can also set parameters such as "bravery level: 8 / 10" and "extroversion level: 6 / 10" using sliders. Users submit this description information to the server by clicking the send button on the terminal interface.

[0101] After receiving input information from the terminal, the server performs a series of specific data processing operations within its internal information processing unit. First, in the text cleaning submodule, the server performs operations such as whitespace normalization, standardization of punctuation, and removal of control characters on the input natural language text to reduce the interference of noisy data on subsequent algorithms. Then, the server uses a natural language processing library to segment the text, breaking long sentences into word sequences. Using a pre-defined dictionary and rules, the server maps these word sequences to high-level abstract features. For example, "brave," "fearless," and "unyielding" are grouped into the "brave" feature dimension, and "value honor" and "protect reputation" are grouped into the "sense of honor" feature dimension, thereby reducing the feature space dimension and improving the stability of subsequent calculations.

[0102] In the attribute extraction submodule, the server extracts features from the input information based on the aforementioned lexical sequences and mapping rules, including personality traits, social roles, activity environments, and behavioral tendencies. The server records these features as key-value data objects; for example, the value of the "occupation" key is set to "knight," the value of the "worldview" key is set to "medieval fantasy," and corresponding fields are reserved for numerical features such as "bravery." This structured data allows the server to precisely control the presentation order and weight of each type of information when constructing the model input, thereby reducing the impact of input noise on the generative artificial intelligence model at the algorithm level.

[0103] In the prompting mechanism, the server generates two types of natural language prompt text based on structured data: one type is system-level prompts, used to constrain the overall behavioral boundaries of the generative AI model; the other type is user-level prompts, used to carry specific character settings and task descriptions. System-level prompts may include the following: "You are a system used to generate character personality prompts. Your task is to output a detailed personality setting prompt based on given character attributes to guide the behavior of a conversational generative AI model. The output language is Simplified Chinese." User-level prompts can be automatically generated based on extracted features, for example: The character attributes are as follows: - Class: Knight - Worldview: Medieval Fantasy - Personality traits: Courageous, values ​​honor, protects the weak - Courage level: 8 / 10 - Extroversion level: 6 / 10 Based on the information above, please generate a character personality prompt statement to guide a generative artificial intelligence model. Requirements: 1. Clearly define the role's values, priorities, and decision-making principles; 2. Describe the character's typical reactions when facing danger, protecting others, or confronting enemies; 3. Specify the character's tone and style (e.g., solemn, polite, and sometimes slightly humorous); 4. The word count should be no less than 200 words. The server concatenates or encapsulates these two types of prompts in a predetermined order into text for model input, enabling the generative AI model to distinguish between global rules and specific task instructions during internal processing. Compared to traditional single-segment prompts, this hierarchical structure creates a clearer hierarchy during the model's attention allocation process, guiding the model to prioritize system-level constraints during the decoding phase, thereby improving the consistency of personality settings.

[0104] The server invokes a generative artificial intelligence model within the prompt generation device. This model can employ a multi-layer self-attention network decoder structure, comprising several stacked Transformer modules. Each Transformer module contains a multi-head self-attention sublayer and a feedforward fully connected sublayer, and integrates information with a normalization layer via residual connections. During the inference phase, the server encodes the model input as a discrete word sequence using text, maps the words to high-dimensional vectors using an embedding matrix, and then performs tensor operations within the multi-layer Transformer. The multi-head attention mechanism at each layer calculates the dot product of the query vector and the key vector, applies scaling and normalization operations, and obtains attention weights for words at different positions, enabling the model to comprehensively consider historical context information to determine the conditional probability of the next output word.

[0105] During the model training phase (which can be completed before system deployment), the server can employ an autoregressive language modeling objective, learning model parameters by maximizing the log-likelihood of the next word in the training corpus. During training, the server uses the cross-entropy loss function to measure the difference between the model's output probability distribution and the true word distribution, calculates gradients using backpropagation, and updates the weights of each layer using optimization algorithms (such as adaptive learning rate optimization methods). The server can also employ data augmentation strategies, such as synonym replacement and sentence structure transformation, to improve the model's generalization ability to diverse inputs. These training details ensure that the generative AI model has good responsiveness to structured prompts during inference.

[0106] During the runtime phase, after receiving the personality prompts output by the model, the server initiates a post-processing unit. This unit first performs content detection, identifying potential violations or inappropriate content through keyword filtering rules and optional auxiliary classification models. Subsequently, the server removes repetitive sentences, uninformative phrases, and obviously redundant introductory statements, and corrects paragraph structure and punctuation, thereby generating standardized text that is easy for both machine and human processing. At this stage, the server can also trim or extend the text to meet preset length constraints, ensuring that the personality prompts are sufficiently detailed yet not overly verbose at the data level, reducing network transmission burden.

[0107] The server returns the post-processed prompt to the terminal via a communication device. The terminal displays the personality prompt as a scrollable text area in its display module, providing controls for "edit," "copy," and "regenerate." Users can add, delete, or modify parts of the sentence to reflect more nuanced creative intentions. For example, a user can add details such as "occasionally uses humor to ease tension" to the original personality prompt. The terminal then sends these edits back to the server, which merges the edits with the original structured data in its update device to form new input text for the model, which is then used to regenerate the personality prompt using a generative AI model.

[0108] To illustrate the specific effects of the system of the present invention in practical use, the following are examples of prompt statements that users can input into the generative artificial intelligence model. These prompt statements can be directly constructed based on the personality prompt statements generated by the server: Example 1: "Personality Setting: You are a knight in a medieval fantasy world, extremely brave, valuing honor and oaths. You habitually speak in a solemn, polite yet firm tone, never backing down in the face of danger, and prioritizing the protection of the innocent and the weak. You abhor cowardice and betrayal, but will give a second chance to those who sincerely repent. In battle, you calmly analyze the terrain and the disparity between yourself and the enemy, but at crucial moments, you will not hesitate to risk your life."

[0109] Task: Based on the above personality traits, write a first-person monologue of approximately 600 words, describing your inner thoughts as you face off against the dragon vanguard on the city wall. Example 2: "Personality Tip: You are a calm and rational urban detective. You speak concisely and directly, avoid unnecessary small talk, are highly sensitive to every detail, and are accustomed to using short questions to dissect others' statements in conversations. You are skeptical of obviously emotional statements and do not easily reveal your emotions."

[0110] Task: Based on the personality clues above, write a dialogue for interrogating a suspect, with no fewer than 20 rounds. In each round, the detective's questioning should progressively narrow down the range of suspects. In terms of technical effectiveness, the server achieves fine-grained control over the data flow of the generative AI model's input and output through the aforementioned structured preprocessing, hierarchical prompt construction, and post-processing mechanisms. Because the input information is converted into structured data, the server can reduce the model's search space during inference without increasing the model's parameter size by reducing irrelevant feature dimensions and explicitly labeling key feature categories, thereby shortening generation time and reducing the probability of erroneous behavior patterns. Furthermore, due to the hierarchical construction of system-level and user-level prompts, the model automatically assigns higher weights to system-level constraints within the self-attention layer, ensuring consistency in character values ​​and behavioral rules across different tasks, thus reducing the setting drift problem common in traditional single-segment prompt methods.

[0111] The server updates the generation conditions using operation and selection history information within the learning device. When a user repeatedly selects a certain style of personality prompt or frequently edits a certain feature, the server can internally maintain a set of preference vectors associated with the user's identifier. This preference vector is used to adjust feature weights when generating new personality prompts, such as increasing the weight of "level of detail" or "emotional expression." Through this parameterized control based on user behavior, the server achieves automatic learning of personalized generation conditions. Compared to simple rule configuration, this more accurately matches the user's actual usage habits, reduces the number of manual parameter adjustments, and improves the overall system interaction efficiency.

[0112] In terms of emotional state regulation, the server receives state information related to the user's emotional state from the terminal, such as the user's active selection of emotion tags like "relaxed," "serious," "comforting," and "encouraging," or the emotion category inferred by the terminal based on the user's interaction behavior. The server then strategically adjusts the generated personality prompts based on this state information within the regulation device. For example, it reduces the intensity of the tone and adds reassuring expressions when the user is in a low mood, and adds challenging and dramatic elements when the user is at their creative peak. By dynamically adjusting the language style and level of detail at the algorithmic level, the server not only improves the user experience but also technically achieves multi-context adaptation for the same character setting, enhancing the reusability of personality prompts in different scenarios.

[0113] The embodiments of this invention can also be extended to various variations. The server can support generative artificial intelligence model combinations of different scales. For example, in resource-constrained scenarios, a small-scale model can be used for initial generation, followed by refinement using a large-scale model; or in online scenarios with strict response time requirements, a quantized or pruned model can be used to reduce inference latency. The server can also offload some preprocessing logic to the terminal, allowing the terminal to perform basic word segmentation and simple feature extraction locally, thereby reducing server load under unstable network or high concurrency conditions and achieving layered distribution of communication load.

[0114] Through the above-described structure, the server in this invention does not merely automate the manual creation of character settings, but rather establishes a dedicated data structure, processing order, and control algorithm within the computer for the task of generating personality prompts, ensuring that the input and output of the generative artificial intelligence model are precisely managed. As a result, the system achieves substantial technical improvements in terms of generation quality, computational efficiency, resource utilization, and communication overhead, realizing an improvement in computer technology itself, rather than simply automating the human creative process.

[0115] use Figure 11 The processing procedure is explained.

[0116] Step 1: In this step, the terminal receives role setting information input by the user.

[0117] Users operate the input controls on the terminal interface, enter text in the character name input box, enter natural language descriptions such as "brave, values ​​honor, protects the weak" in the personality description text box, select options such as "knight" in the character type selection box, and set numerical parameters such as "bravery level: 8 / 10" and "extroversion level: 6 / 10" through the slider.

[0118] The input to the terminal consists of raw text strings and option selections provided by the user through input devices such as keyboards and touchscreens.

[0119] The terminal's input processing includes: reading the current values ​​of each interface control, combining text fields, selection fields, and numeric fields into a set of key-value pair data structures, and performing basic null and format checks, such as confirming that required fields are not empty.

[0120] The terminal's output in this step is: the raw input object containing character attribute and behavioral information, such as structured data including fields like "personality description," "occupation type," "worldview," and "numerical parameters," which is prepared to be sent to the server in the next step.

[0121] Step 2: In this step, the terminal packages the user-input data and sends it to the server.

[0122] The terminal input is the original input object generated in step 1.

[0123] The terminal performs the following data processing and communication operations: It uses its internal network library to serialize the raw input object into a text format (e.g., a JSON string), constructs a network request message in memory containing the target address, HTTP headers, and a request body, and sends this message to the server via the network interface. Simultaneously, the terminal initiates an asynchronous wait process locally to receive the server's response.

[0124] The terminal's output in this step is a request data packet transmitted over the network, the logical content of which is "user role setting input information". This data packet becomes the input for the server's next step.

[0125] Step 3: In this step, the server receives and parses the input information from the terminal.

[0126] The server's input is a request data packet sent by the terminal over the network, which contains serialized role setting data.

[0127] The server's specific processing includes: receiving data packets through its network interface and passing the messages to the application layer program; using its parsing module to decapsulate the request body from the network protocol and parse the JSON string into an internally operable data structure (such as a dictionary or object); and then performing existence, type, and length checks on the fields, for example, ensuring that "personality description" is a string with a preset length, and that "courage level" is a numerical value within the range of 0 to 10. For missing or invalid fields, the server can generate error messages.

[0128] The server's output in this step is: the verified role input data object, which will serve as the input for subsequent preprocessing steps.

[0129] Step 4: In this step, the server performs text cleaning and normalization preprocessing on the character input data.

[0130] The server's input is the role input data object obtained in step 3, which includes multiple natural language text segments and several numerical parameters.

[0131] The data processing operations performed by the server include: in the text cleaning submodule, the server removes redundant spaces, tabs, and invisible control characters, unifies full-width punctuation to half-width or vice versa (according to a predetermined strategy), and standardizes line breaks; in the language standardization module, the server can merge synonyms, for example, marking "heroic" and "fearless" as phrases of the same kind as "brave," providing a unified word form for subsequent feature extraction.

[0132] Through the above processing, the server transforms the non-normalized text into a more regular set of strings, reducing noise and redundancy.

[0133] The server output in this step is a set of cleaned and normalized text fields and retained numerical fields, which will serve as input for the next step of feature extraction and structured processing.

[0134] Step 5: In this step, the server segments the normalized text into words and extracts character features to construct structured data.

[0135] The server's input consists of the normalized text and numeric fields output from step 4.

[0136] The specific data processing performed by the server includes: calling a natural language processing library to perform Chinese word segmentation on fields such as "personality description" and generating word sequences; using a pre-built dictionary and rules, mapping words to high-level feature categories, such as mapping "brave," "fearless," and "unyielding" to "brave feature," and mapping "values ​​honor" and "defends reputation" to "sense of honor feature," and counting the frequency of occurrence or assigning scores; the server also directly assigns numerical parameters to the corresponding feature fields, such as recording the bravery level of 8 / 10 as the "brave intensity" attribute.

[0137] The server then constructs structured data objects, organizing "personality traits," "occupation," "worldview," "behavioral tendencies," and "numerical characteristics" into key-value pairs or table structures, so that each feature category corresponds to a set of standardized values.

[0138] The server's output in this step is: a set of structured character feature data, containing abstracted attributes and behavioral elements. This structured data will be used as input for the prompt statement construction stage.

[0139] Step 6: In this step, the server constructs system-level and user-level prompts and generates text for model input.

[0140] The server's input is the structured feature data obtained in step 5.

[0141] The data processing operations performed by the server include: In the prompt construction module, based on predefined text templates and structured features, inserting different feature fields into corresponding text placeholders. For example, the server fills in "Occupation: Knight," "Worldview: Medieval Fantasy," and "Personality Traits: Brave, Honor-conscious, Protects the Weak" into a descriptive text to generate a user-level prompt statement; simultaneously, the server generates a fixed system-level prompt statement to explain the generation task and output requirements. The server places the system-level prompt statement at the beginning according to a predetermined order, places the user-level prompt statement in the secondary part, and inserts task requirements, such as "at least 200 words," etc.

[0142] The server concatenates multiple text segments, inserts line breaks, and adds labels (such as using specific prefixes to indicate system constraints) at the string level to construct the complete text for model input.

[0143] The server's output in this step is: the combined model input text, which will be fed into the generative artificial intelligence model as inference input.

[0144] Step 7: In this step, the server encodes the model input into text and inputs it into a generative artificial intelligence model for reasoning, generating initial personality prompts.

[0145] The server's input is the model input text output in step 6.

[0146] The data computation performed by the server includes: first, using a word segmentation and encoding module to convert the model input text into a sequence of tokens, and then mapping each token to a high-dimensional vector representation; subsequently, the vector sequence is input into a multi-layer Transformer decoder structure, with each layer performing multi-head self-attention operations and feedforward network operations: in the self-attention calculation, the server calculates the correlation between the query vector and the key vector through matrix multiplication and dot product operations, and obtains attention weights through scaling and normalization, and then uses these weights to perform a weighted summation of the value vectors to obtain a context-aware representation. Through multi-layer stacking, the model gradually aggregates long-term dependency information.

[0147] During the decoding process, the server selects the highest probability lexical from the vocabulary or selects lexical according to the sampling strategy based on the current context vector and the output lexical probability distribution, and gradually generates each word in the personality prompt statement.

[0148] The server's output in this step is a raw personality cue text generated by a generative artificial intelligence model. This text typically includes a description of the person's values, behavioral principles, response methods, and language style.

[0149] Step 8: In this step, the server performs post-processing on the generated personality prompts, including content inspection and formatting.

[0150] The server's input is the original personality prompt text output in step 7.

[0151] The data processing operations performed by the server include: In the content detection module, the server checks whether the text contains prohibited words, sensitive content, or obviously non-compliant sentences through keyword matching and optional classifiers; if non-compliant paragraphs are found, the server can mark or delete them; In the duplicate detection module, the server identifies and deletes obviously duplicate or redundant sentences by comparing sentence vectors or text hashes; In the formatting module, the server standardizes punctuation, paragraph division, and line breaks to make the text structure clear.

[0152] The server can also trim or expand the text according to preset upper and lower length limits, such as by regenerating some paragraphs or adding summary sentences to make the text meet the requirement of "no less than 200 words".

[0153] The server's output in this step is: a filtered and formatted standardized personality suggestion text, ready to be sent to the terminal for display.

[0154] Step 9: In this step, the server sends the standardized personality prompt statement to the terminal and records the relevant generation information.

[0155] The server's input consists of the standardized personality prompt statements output in step 8, along with the corresponding structured feature data and user identifiers.

[0156] The server performs the following operations: In the communication module, it encapsulates the personality prompt statement and metadata (such as generation time and associated personality attribute summary) into a response data structure, constructs a response message in the network protocol stack, and sends it to the terminal through the network interface; In the storage module, it writes the personality prompt statement, the corresponding input information, generation conditions, and user identifier into the storage device for subsequent statistical and personalized learning purposes.

[0157] The server's output in this step is: a response data packet sent to the terminal (the logical content of which is the final personality prompt statement) and historical data written to the storage device. This data becomes the input for the terminal display and the server learning module.

[0158] Step 10: In this step, the terminal receives and displays personality prompts, and also receives editing or regeneration instructions from the user.

[0159] The terminal input is the response data packet returned by the server in step 9.

[0160] The data processing performed by the terminal includes: parsing response messages in the network module to extract the personality prompt text and metadata; displaying the personality prompt in a scrollable text area in the interface module, and presenting "Edit," "Copy," and "Regenerate" buttons on the interface. If the user modifies the text, the terminal will combine the edited text with the original metadata to form a new edit information object; if the user clicks "Regenerate," the terminal generates control data containing regeneration instructions.

[0161] The terminal's output in this step consists of two parts: a personality prompt statement interface displayed to the user and editing information or regeneration instruction data sent to the server, which are used as input for the server to update the generation conditions and regenerate the prompt statement.

[0162] Step 11: In this step, the server updates the generation conditions based on the user's edited information and operation history, and may trigger a new round of personality prompt statement generation.

[0163] The server's input consists of the editing information and regeneration instructions sent by the terminal in step 10, as well as the existing operation history information and selection history information in the storage device.

[0164] The data computation performed by the server includes: statistically analyzing the frequency and direction of user editing behavior for different features in the learning device. For example, if a user frequently enhances descriptions related to "emotional expression", the server will increase the weight of the corresponding feature in the user's preference vector accordingly; the server will merge the new editing information with the original structured features according to the regeneration instructions, construct updated structured data, and re-enter the prompt construction and model inference process.

[0165] In this way, the server gradually adjusts the generation condition parameters (such as feature weights, text detail, tone intensity, etc.) to better match user preferences in subsequent generation.

[0166] The server's output in this step is: updated user-specific generation conditions (such as preference vectors and rule configurations) and a new round of generated personality prompts. These outputs are then used by the terminal for display and further interaction, thus forming a continuously optimized generative personality prompt system.

[0167] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0168] While existing content generation and distribution technologies include solutions that utilize generative artificial intelligence models to automatically generate text content, they still have significant shortcomings in the following aspects, thus limiting the improvement of the overall processing power of computer systems and user interaction experience.

[0169] First, existing systems typically treat generative AI models simply as "black box text generators," generating stories and other content based solely on fixed prompts input in a one-time manner, lacking the ability to manage and dynamically update the prompts themselves in a structured way. Computer systems struggle to finely control the prompts, especially in constructing programmable prompt generation processes around specific object attributes (such as character traits and behavioral styles), resulting in poor controllability and consistency of the generated results.

[0170] Secondly, regarding terminal-server collaborative processing, traditional systems often adopt a model where the terminal directly inputs complete prompts, and the server generates the content in one go. The server side lacks intermediate state management and reuse mechanisms for prompt statements, and it cannot tightly link the automatic generation of prompt statements, user editing, and story generation into an integrated processing pipeline. This loosely coupled architecture not only increases the user's operational burden but also makes it difficult for the server to optimize the subsequent generation process based on the fine-grained interactions of the user on the terminal, resulting in low utilization efficiency of computing resources.

[0171] Third, regarding the dynamic generation and real-time presentation of long text content such as stories, existing technologies often adopt a "return everything after generation" approach, lacking a segmented transmission and sequential display control mechanism for streaming output between the server and the terminal. As a result, when the output scale of the generative AI model is large, it not only increases network transmission latency but also causes problems such as large one-time rendering overhead and sluggish interface response on the terminal side, making it difficult to realize an interactive content consumption model of "generating and experiencing simultaneously".

[0172] Fourth, to address the need for diverse story development based on different object attributes, traditional solutions often rely on users manually changing character settings, re-entering prompts, and re-initiating generation requests at the application layer. This lacks a programmable, repeatable attribute-driven generation control logic for computer systems. The server side cannot automatically control the cyclical process of "prompt generation—user editing—story generation—real-time distribution" based on attribute information uploaded from the terminal. This results in high overall process complexity and maintenance costs when the system is expanded to multi-character, multi-worldview scenarios.

[0173] Therefore, a new system architecture and processing method are needed to enable servers to automatically construct and update prompts for generative artificial intelligence models based on object attribute information, and to integrate prompt generation, user editing, dynamic story generation, and real-time distribution in collaboration with the terminal, thereby improving the computer level: (1) The structured generation and controllability of prompt statements; (2) Data processing pipeline between terminal, server and generative artificial intelligence model; (3) Mechanism for streaming generation and real-time presentation of long text content; (4) The ability to automatically control the repeated generation of different story developments based on attribute information.

[0174] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.

[0175] In this invention, the server includes a processing unit for automatically generating prompt statements for input into a generative artificial intelligence model based on object attribute information from a terminal; a management unit for sending the generated prompt statements to the terminal and receiving updated prompt statements edited or added by the user via the terminal; a generation unit for dynamically generating story information by calling the generative artificial intelligence model based on the updated prompt statements and the object attribute information; a distribution unit for sending the generated story information to the terminal in segments in real-time or streaming mode and controlling the terminal to display the story information sequentially; and a control unit for repeatedly executing prompt statement generation and story generation processing based on new attribute information from the terminal to provide diverse story developments. This creates a closed-loop processing pipeline within the computer system, from attribute-driven automatic prompt statement generation to user interactive editing, and then to dynamic story generation and real-time distribution based on the generative artificial intelligence model. This improves the programmability and controllability of the prompt statement and story generation process, reduces the terminal input burden, enhances the consistency between model inference results and object attributes, and reduces network and rendering latency through a streaming output mechanism. It enables real-time generation and experience of diverse stories based on object attributes, thereby improving the overall performance of the computer technology solution and the quality of user interaction.

[0176] "System" refers to an integrated technical solution consisting of at least one information processing device, at least one terminal, and programs and communication interfaces running on it, used to realize the generation, updating, story generation and distribution of prompt statements, etc.

[0177] "Information processing device" refers to a data processing unit with computing and storage resources, used to execute programs and process attribute information, prompts, and story information from a terminal. It can be a server, a cloud computing node, or other computing device.

[0178] "Terminal" refers to an electronic device operated by a user and interacting with an information processing device through a communication network, including but not limited to mobile terminals, head-mounted displays, and computing terminals, used to display prompts and story information and receive user input.

[0179] "Object" refers to the target entity that is abstractly represented in the system as the basis for generating prompts and story information, including but not limited to fictional characters, user roles, virtual agents, or other entities with attribute characteristics.

[0180] "Object attribute information" refers to structured or semi-structured data used to describe the characteristics of an object, including but not limited to personality traits, behavioral tendencies, background settings, worldview information, etc., which can be sent by the terminal to the information processing device as input for generating prompt statements and story information.

[0181] "Generative AI models" refer to AI models that automatically generate text or other content based on input data. They are usually based on deep learning structures and acquire generative capabilities through training on large-scale data. They are used to generate prompts and story information.

[0182] "Prompt statements" refer to textual information used as input to generative artificial intelligence models. They include constraints on the generation goal, context settings, and generation requirements, which guide the generative artificial intelligence model to output results that conform to the desired style and content.

[0183] "Personality cue" refers to a form of cue statement, which is textual information used to instruct an object on its behavior and response style in a specific situation. It is used to fix or reinforce the object's personality-related settings in generative artificial intelligence models.

[0184] "Story information" refers to text content generated by generative artificial intelligence models based on prompts and object attribute information. It usually presents the behavior, dialogue and plot development of objects in a narrative form, and is intended for users to read or experience on the terminal.

[0185] A “generation unit” refers to a functional module in an information processing device that is used to call a generative artificial intelligence model and generate story information or other output content based on prompts and object attribute information. It can be implemented through software, hardware or a combination thereof.

[0186] The "management unit" refers to a functional module in an information processing device used to send prompt statements to the terminal, receive updated prompt statements uploaded by the terminal that have been edited or added by the user, and manage the version and status of the prompt statements.

[0187] The “distribution unit” refers to a functional module in an information processing device that is used to send the generated story information to the terminal in real time or streaming mode, and to segment, cache and control the transmission of the story information when necessary.

[0188] "Control unit" refers to a functional module in an information processing device used to control the execution order and repetition number of the prompt statement generation process and the story generation process based on new attribute information from the terminal, thereby realizing the scheduling and management of different story developments.

[0189] "Real-time mode" refers to a processing method in which the information processing device immediately sends the story information or a portion thereof to the terminal during or shortly after the story information is generated by the generative artificial intelligence model, so that the terminal can display it in a way that minimizes the perceived latency for the user.

[0190] "Streaming" refers to a transmission and display method in which the information processing device generates story information and sends the generated results to the terminal in segments or sequences, so that the terminal can display the story content one by one as the data arrives, without waiting for the complete story to be generated.

[0191] "Editing or adding" refers to the user's actions of modifying the prompt statement on the terminal interface, including adding, deleting, or modifying existing text content as well as inputting new text content, thereby changing or expanding the meaning and details of the prompt statement.

[0192] "Different story developments" refers to multiple story versions generated by a generative artificial intelligence model that differ in plot development, character behavior, dialogue content, or ending direction when object attribute information or prompts change.

[0193] In various embodiments of the present invention, the server operates as an information processing device on a computer device with high-performance computing resources, and the terminal operates as a user interaction device on a mobile computing device or a wearable display device. The user interacts with the server through the terminal to realize the generation, editing and dynamic generation and distribution of prompts based on a generative artificial intelligence model.

[0194] In one embodiment, the server utilizes a combination of a general-purpose processor and a graphics processing unit (GPU) to form a hardware platform. The server may include hardware resources such as a multi-core central processing unit (CPU), a graphics processing unit (GPU), main memory, non-volatile memory, and a network interface. In a specific configuration, the server may be based on a general-purpose processor with multi-core processing capabilities and a GPU with massively parallel floating-point operation capabilities, running an operating system, and deploying deep learning frameworks (e.g., frameworks based on tensor operation libraries), generative artificial intelligence model inference frameworks (e.g., inference frameworks based on transformer architectures), network application frameworks, and a database management system. Within this hardware and software environment, the server executes program modules, converting data from the terminal into input prompts adapted to the generative artificial intelligence model, and post-processing and streaming the story information output by the model.

[0195] In one embodiment, the terminal is a handheld mobile terminal. The terminal includes hardware components such as a processor, graphics processing unit, memory, display device, input device, and communication module. The terminal runs an application on a mobile operating system. This application is responsible for presenting the user with a character attribute selection interface, a human-computer interaction interface, a text editing interface, and a story reading interface, and exchanges data with a server through a network interface. In another embodiment, the terminal can be a head-mounted display device. In this case, the terminal uses a 3D rendering engine to map story information into text panels or subtitles in a virtual scene, thereby presenting the generated content in an environment with a more realistic immersive experience.

[0196] Users select object attribute information and edit prompts via the terminal. In the character selection interface, users choose basic attributes of the target object, such as "name," "personality traits," "background setting," and "worldview," and can add additional attributes through multi-selection or text input. After receiving the prompts generated by the server, users modify, supplement, or delete the prompts using the text editing component on the terminal to create more predictable personality prompts, thereby indirectly controlling the behavior patterns of the generative artificial intelligence model.

[0197] In one implementation, the server uses a generative artificial intelligence model employing a transformer neural network architecture. This model is deployed as a text generation module on the graphics processing unit. The model includes multiple self-attention sublayers, a multi-head attention mechanism, feedforward network sublayers, and normalization sublayers. During model inference, the server converts prompts into word or phrase-level labeled sequences, embeds these sequences as high-dimensional vectors, and inputs them into the model's encoder-decoder or autoregressive structure. Within each layer, the model uses attention weights to calculate the correlation between positions in the input sequence, combines positional encoding and nonlinear transformations to generate output vectors, and then maps these output vectors to a probability distribution over the vocabulary using linear transformations and a soft maximum function, thereby generating a labeled sequence of story text in sequence.

[0198] During the model training phase, the server uses a large dataset containing character descriptions and story text pairs to pre-train and fine-tune the generative AI model. The server employs cross-entropy loss as the error function during training to maximize the log-likelihood of the next true label given a sequence of context labels. When updating network parameters, the server uses stochastic gradient descent-like optimization algorithms, such as adaptive moment estimation or its variants, to solve for the loss gradient on mini-batch training samples and update the model weights, thereby enabling the model to learn appropriate response patterns to prompts. The server can also use data augmentation techniques during training, such as paraphrasing character descriptions and varying sentence structures, to improve the model's robustness to different expressions. Through this training method, the server enables the model to more accurately generate consistent and stylistically stable story information based on personality prompts and object attribute information during the inference phase, thereby improving the controllability and accuracy of the generated results.

[0199] The server employs a multi-stage data structure and algorithm in its prompt generation process. First, the server organizes the object attribute information from the terminal into a key-value structure, such as storing it in an in-memory association table as a "attribute name - attribute value list". The server then selects an appropriate language template from a predefined set of prompt templates, each containing placeholders and control instructions. When filling the templates, the server inserts attribute values ​​into the corresponding placeholder positions according to fixed rules, and adds logical conditional statements based on attribute combinations when necessary. For example, when the object's personality trait includes "humor," it adds constraint descriptions related to humorous dialogue. This rule-based template filling algorithm makes the input prompts for the generative AI model more structurally stable and semantically accurate. This structured input helps the model reduce the generation of irrelevant content, decrease inference errors, and reduce the generation of invalid tokens, thereby improving the model's inference efficiency.

[0200] In one example, the server converts object attribute information into prompt statements in the following form to generate personality prompts: Your task is to generate a "personality prompt statement" based on the following character attributes: Character Name: Brave Knight Core personality traits: Courageous, upright, humorous Worldview: Medieval fantasy world Require: 1. Describe the character's typical behavior and reactions in the second person when in danger, talking to companions, or facing help from the weak; 2. The tone should be "brave but not reckless, upright yet humorous"; 3. This text will serve as character design for subsequent story generation; please be as specific as possible. 4. Keep the word count around 200 words.

[0201] In another example, the server combines the user-edited personality prompt with the story task into the following story generation prompt: Character settings As a brave knight, you always stand at the forefront of battle. You prioritize protecting your companions and the innocent, even at your own risk. You speak directly, yet you use humor to ease tension. When facing formidable enemies, you never back down, but calmly assess the situation and seek opportunities for victory while upholding chivalry.

[0202] Story Quest Based on the above character setting, please write the opening of the story, set in a medieval fantasy world, describing the knight's first encounter with a dragon attack on the royal city. Requirements: 1. Narrated in the third person; 2. It must include environmental descriptions (such as the sky, city walls, flames, etc.); 3. It should include at least three rounds of dialogue, showcasing the knight's humor and bravery; 4. Keep the word count around 800 words.

[0203] In the aforementioned processing, the server not only performs simple concatenation of natural language text, but also encodes role settings, task requirements, and text formatting requirements into prompts through specific structured rules and control instructions. This approach enables generative AI models to distinguish between "role setting areas," "task instruction areas," and "formatting requirement areas" within their internal attention mechanisms. This allows them to focus on key sentences and segments during attention calculations, improving the relevance of generated content and reducing irrelevant and redundant content. Through this region-specific and weighted prompt design, the server indirectly constrains the model's internal attention distribution, technically improving the accuracy and efficiency of model inference.

[0204] The server employs segmented encoding and streaming technology in the distribution of story information after generation. After obtaining the marked sequence generated by the model, the server does not assemble it into a complete text all at once before sending it. Instead, it segments the text into multiple small fragments according to paragraph boundaries or fixed lengths, and encodes them sequentially into small message units. At the network layer, the server uses communication protocols that support long connections or push notifications to send these message units to the terminal in near real-time. During transmission, the server dynamically adjusts the length and transmission interval of each fragment based on current network bandwidth and terminal feedback to balance latency and stability. This streaming mechanism reduces the data peak of a single transmission, helping to lower the probability of network congestion and the pressure on the terminal for simultaneous rendering, thus achieving a "generate and display simultaneously" user experience. It also optimizes bandwidth utilization and buffer management at the computer system level.

[0205] When receiving story information fragments, the terminal decodes and buffers each fragment, then appends them sequentially to the text display area. At the interface rendering layer, the terminal uses an incremental layout algorithm, only rearranging and drawing areas related to newly added content, thus avoiding a full redraw of the entire interface and reducing the load on the graphics processing unit and central processing unit. At the application logic layer, the terminal controls scrolling and highlighting effects based on the order and time intervals of fragment arrival, providing users with a continuous and smooth reading experience. Through this incremental rendering strategy, the terminal achieves efficient presentation of long texts with limited hardware resources.

[0206] The server uses a specific data structure for managing and versioning prompt updates. It associates each object instance with its corresponding personality prompt version using identifiers, employing a two-tiered storage structure including "current version" and "historical versions." When the server receives an update prompt from the terminal, it calculates the differences between the updated content and the previous version, generates a change record, and assigns a version number and timestamp to the new version. This version management mechanism allows the server to quickly roll back to a historical version in case of subsequent generation failures or user undoing operations. Furthermore, it can analyze user editing behavior characteristics based on multi-version data to optimize subsequent automatic generation strategies, thereby gradually reducing the amount of manual modification work required by users at the algorithmic level and achieving adaptive optimization of prompt quality by the computer system.

[0207] In one implementation, the server employs a hybrid strategy combining rule-driven and model-generated approaches, distinguishing itself from traditional methods that rely entirely on manual input or model-driven free generation. In the initial stage of prompt generation, the server uses deterministic templates and rule controls to map object attributes to descriptions with fixed structures, avoiding excessive instability caused by model randomness. In subsequent stages of personality detail refinement and story content creation, it primarily relies on generative AI models to expand upon the constraints outlined earlier. This unconventional hybrid process of "rule constraints + model generation" ensures the system maintains high controllability while preserving the expressive power of the generative AI model, technically improving the balance between consistency and diversity in generated content.

[0208] The server employs a modular design for internal data flow. It separates the attribute processing module, prompt statement construction module, model interface module, result post-processing module, and distribution module, connecting them via message queues or function calls. The server transmits structured data objects between modules, each containing necessary metadata such as object identifier, version number, and content type. This modular and structured data design facilitates the replacement of specific generative AI models, optimization of transmission protocols, or adjustment of prompt templates in different implementations without affecting the overall workflow, thereby improving the system's scalability and maintainability.

[0209] In various embodiments of this invention, the server, through a combination of the aforementioned specific data structures, algorithmic processes, and model structures, achieves a technical improvement over the traditional method of "humans manually writing prompt text + one-time call to the text generation interface." By automatically constructing structured prompt statements and introducing personalized prompt version management and streaming distribution mechanisms, the server achieves fine-grained control and resource optimization of the large-scale text generation process within the computer. This results in improved generation accuracy, reduced network and rendering latency, increased computing resource utilization, and enhanced real-time user interaction, rather than simply automating human writing.

[0210] use Figure 12 The processing procedure is explained.

[0211] Step 1: Users select object property information on the terminal.

[0212] Input: Pre-displayed candidate information for objects on the terminal (such as multiple character names, personality tags, worldview options, etc.).

[0213] Output: Object attribute information processed by the terminal (such as structured data containing fields such as "name, personality traits, background setting, worldview").

[0214] The terminal displays a list of roles and attribute options through a graphical user interface. Users select one or more roles via a touchscreen or other input device, and then check personality tags and input background descriptions for the selected roles. After detecting user confirmation, the terminal aggregates these scattered input items into an internal data structure to form object attribute information, and temporarily caches it in local memory, awaiting transmission to the server.

[0215] Step 2: The terminal sends object attribute information to the server.

[0216] Input: Object attribute information maintained internally by the terminal.

[0217] Output: The object attribute information request message sent to the server over the network, and the object attribute data received and parsed by the server.

[0218] The terminal serializes the object attribute information into a network transmission format and sends the request to the interface address specified by the server by invoking a secure transmission protocol through the communication module. After receiving the request, the server parses the request message through the network stack and application framework, extracts the object attribute data from the message body, and converts it into a server-side memory data structure for subsequent processing.

[0219] Step 3: The server constructs prompt statements and generates template input based on object attribute information.

[0220] Input: Object attribute information in server-side memory (including object name, personality traits, worldview, background description, etc.).

[0221] Output: Template text for generating prompts for generative artificial intelligence models.

[0222] After receiving the object attribute information, the server selects a template matching the object type and application scenario from a stored set of prompt templates, and fills the placeholder positions in the template with the object attribute information. During the filling process, the server concatenates the personality trait list, controls the length of the background description, and normalizes the worldview field. Then, it combines these elements to generate a structured text that clearly explains the model task and format requirements, serving as input prompts for the generative artificial intelligence model.

[0223] Step 4: The server calls a generative artificial intelligence model to generate personality prompts.

[0224] Input: Template text generated by the prompt statement constructed by the server in step 3.

[0225] Output: Personality prompt text output by the generative artificial intelligence model.

[0226] The server inputs template text into a generative AI model deployed on a graphics processing unit. The model converts the text into a tokenized sequence using word segmentation or sub-word fragmentation mechanisms, embedding this sequence into a high-dimensional vector space before feeding it into the model's transformer network. Internally, the model performs multi-layered transformations on the vector sequence using self-attention and feedforward networks. At each time step, the server calculates the probability distribution of the next token based on the output vector and samples it, progressively generating a tokenized sequence of personality tips. After generation, the server reverse-maps the tokenized sequence back into natural language text and uses this text as a candidate result for personality tips.

[0227] Step 5: The server performs post-processing on the personality prompts and sends them to the terminal.

[0228] Input: The original personality prompt text output by the generative artificial intelligence model.

[0229] Output: Cleaned and length-controlled personality prompt text, and a response message sent to the terminal.

[0230] The server performs text cleaning on the raw output, removing irrelevant opening phrases, repeated paragraphs, or extra spaces, and truncates or segments the text to meet system-defined limits. The server can also perform sensitive content detection and simple rule filtering to ensure the content remains within agreed-upon limits. Subsequently, the server encapsulates the processed text along with metadata such as object identifiers into response data and transmits it back to the terminal via network protocols. Upon receiving the response, the terminal parses the personality prompt text from the message and caches it for subsequent display and editing.

[0231] Step 6: The terminal displays personality prompts and allows users to edit them.

[0232] Input: The personality prompt text received and parsed from the server.

[0233] Output: The final personality prompt text edited by the user.

[0234] The terminal displays personality prompts on its interface and provides text editing controls, allowing users to add, modify, or delete sentences. Users can fine-tune the personality prompts via input devices, such as adding specific behavioral styles, adjusting tone, or restricting certain behaviors. The terminal updates its internal text cache in real time with each user input and, after the user confirms completion of editing, prepares the final version of the personality prompt as output to the server.

[0235] Step 7: The terminal sends the final personality prompt to the server.

[0236] Input: User-edited personality prompt text cached within the terminal and associated object identifier.

[0237] Output: The prompt update request message sent to the server, and the final personality prompt data received and stored by the server.

[0238] The terminal packages the edited text and object identifier together as request data and sends it to the server via the network interface through the communication module. Upon receiving the request, the server parses the text content and identifier, writes the new prompt text to the storage system, updates the corresponding object's personality prompt version record, and retains the latest prompt in memory for subsequent story generation.

[0239] Step 8: The server generates story prompts based on the final personality prompts.

[0240] Input: The final personality prompt text and object attribute information stored on the server.

[0241] Output: Story-generating prompt text for generative artificial intelligence models.

[0242] After obtaining the latest personality cues, the server combines them with the character's worldview and the current story task settings (e.g., "first confrontation with the enemy," "investigating the crime scene," etc.). Through string concatenation and template filling, the server places the personality settings in the "Character Setting" section and the story task, narrative perspective, length requirements, and dialogue quantity in the "Task Requirements" section, forming a clearly structured story generation prompt. During the construction process, the server adds paragraph labels and explanatory text to each part so that the generative AI model can distinguish different functional areas within its attention mechanism.

[0243] Step 9: The server calls a generative artificial intelligence model to generate story information.

[0244] Input: The story-generating prompt text constructed in step 8.

[0245] Output: Story information text output by the generative artificial intelligence model.

[0246] The server feeds story-generating prompts into the generative AI model, using the same transformer architecture and inference process as the personality prompt generation model, but with a larger maximum output length parameter to allow for the generation of longer story texts. During inference, the server controls text diversity based on pre-defined parameters such as temperature and sampling strategies, and monitors the number of generated tags, stopping generation when it approaches the target length or reaches the termination condition. The server decodes the tag sequence output by the model into natural language story text, outputting it as raw story information.

[0247] Step 10: The server segments the story information and sends it to the terminal in a streaming manner.

[0248] Input: The complete story information text output by the generative artificial intelligence model.

[0249] Output: Segmented story text fragments and message sequences streamed to the terminal over the network.

[0250] The server divides the story text into multiple logical segments based on sentence boundaries or paragraph markers, and adds a sequence number and timestamp to each segment. The server selects an appropriate communication method in the network transmission module and sends these segments to the terminal one by one in near real-time. During transmission, the server dynamically adjusts the segment size and transmission rate based on network conditions to avoid network congestion, and internally maintains a transmission queue and retransmission strategy to ensure that the segments arrive at the terminal reliably in order.

[0251] Step 11: The terminal receives story clips and displays them sequentially.

[0252] Input: A story text fragment message streamed by the server.

[0253] Output: The story content presented sequentially in the terminal display interface.

[0254] After receiving each story fragment, the terminal decodes the message and appends the text content to the local story cache. At the user interface layer, the terminal appends the newly arrived fragment to the bottom of the reading area, while simultaneously performing local layout and rendering operations on the newly added area to avoid a full-screen redraw. The terminal can automatically scroll or highlight new text based on the frequency of fragment arrivals to help users follow the story's progression. This segment-by-segment display method allows users to begin reading existing content while the server is still generating subsequent fragments.

[0255] Step 12: Users can read stories on the terminal and reselect object attributes as needed to obtain different story developments.

[0256] Input: The story content currently displayed on the terminal and the available object switching or attribute adjustment options.

[0257] Output: Information triggered by the user's new object property selection and / or new edit request.

[0258] While reading a story, users can trigger actions such as "reselect character" or "adjust personality traits" on their devices if they wish to experience different personality settings or plot developments. Upon receiving this action, the device guides the user back to the attribute selection interface or prompt editing interface, generating new object attribute information or a new prompt editing request. The device then sends these new inputs back to the server, triggering a cyclical execution of the aforementioned steps, allowing the system to generate new story developments based on the new attributes and prompts. In this way, users can experience diverse stories multiple times using the same system, while the server, through repeated data processing and calculations, achieves an attribute-driven dynamic content generation process.

[0259] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.

[0260] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."

[0261] With the widespread application of generative artificial intelligence models in text generation, image generation, and other fields, users often need to write high-quality prompts as input to the model in order to obtain the desired generated results. However, existing technologies suffer from the following problems: (1) In the traditional human-computer interaction process, users often only input one-time, rough natural language instructions into the generative artificial intelligence model. There is a lack of systematic generation and optimization mechanism for the prompt statements themselves, which makes the quality of the prompt statements highly dependent on the user's experience and language expression ability. In this process, the computer system only acts as a passive execution tool and fails to effectively control the generation process at the prompt statement design level.

[0262] (2) Although some existing systems can automatically generate preliminary prompt text based on user input, they usually only provide a one-way "input → generate" function. They lack a multi-round interactive optimization process around the prompt statement and cannot record and utilize the user's multiple edits and the model's evaluation results in a structured manner. This results in insufficient computer modeling ability for user goals and poor consistency and controllability of the generated results.

[0263] (3) In the prior art, most generation systems do not manage the versions of prompt statements in a fine manner, and lack the mechanism for saving, backtracking and comparing different versions of prompt statements. Users find it difficult to effectively weigh multiple candidate solutions and accumulate reusable prompt statement assets in long-term use. The computer system cannot mine and learn from these historical data, thus limiting the improvement of the overall system performance and user creation efficiency.

[0264] (4) Typical natural language interfaces only return the generated result itself, without providing structured feedback information on the prompt statement, such as advantages, improvement points and revision schemes. Users find it difficult to understand the causal relationship between the prompt statement and the generated result. The computer also cannot explicitly expose the generative artificial intelligence model's "evaluation ability" of the prompt statement to the user and incorporate it into the interaction loop. As a result, the prompt statement optimization process lacks clear computational goals and visual guidance signals.

[0265] Therefore, the technical challenge to be addressed by this invention is how to introduce a multi-stage interactive generation and optimization mechanism for prompt statements into a computer system, enabling the server to: automatically generate initial prompt statement schemes using generative artificial intelligence models; perform machine evaluation and feedback generation on user-customized prompt statements; and support users to conduct multiple rounds of iterative design by combining version management and historical backtracking, thereby improving the efficiency and quality of the input construction of generative artificial intelligence models at the system level.

[0266] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.

[0267] In this invention, the server includes a functional module for performing natural language processing parsing on information obtained from the user using a generative artificial intelligence model, and automatically generating an initial scheme of prompt statements as input to the generative artificial intelligence model based on the parsing results; a functional module for presenting the initial scheme of prompt statements on the user interface of a communication terminal with display and input functions and receiving user additions, deletions, or modifications to the prompt statements to achieve user customization; a functional module for storing the user-customized prompt statements in a storage device in association with version information, and inputting the customized prompt statements into a generative artificial intelligence model or a generative artificial intelligence model for evaluation to generate feedback information including advantages, improvement points, and revision schemes; and a functional module for presenting the feedback information on the user interface and enabling the user to repeatedly edit the prompt statements based on the feedback information to refine the prompt statements through multi-stage interaction, and outputting the final prompt statements as input data to an external generative artificial intelligence model or an external information processing device. This allows for a closed-loop optimization process oriented towards prompts within the computer system. The server not only undertakes the computational task of generating content, but also actively participates in the automatic generation, version management, and iterative optimization based on model evaluation of prompts. This improves the automation and quality stability of the input construction of generative artificial intelligence models, enhances the human-computer collaborative creation process, and improves the technical performance of the entire generation system in terms of accuracy, controllability, and reusability.

[0268] "Generative artificial intelligence models" refer to data processing models that use machine learning algorithms to train on large-scale data and can automatically generate corresponding text, images or other content based on input text or other data.

[0269] "Prompt statements" refer to natural language text used as input to generative artificial intelligence models, which are instructions to the generative artificial intelligence models to perform specific generative tasks or to output according to specific style and content requirements.

[0270] The "initial prompt statement scheme" refers to the prompt statement version that the server automatically generates based on user input information using natural language processing and generative artificial intelligence models, and which has not yet been edited by the user.

[0271] "User interface" refers to the collection of display screens and input controls presented on a communication terminal for information interaction with the user, including graphic elements such as text boxes, buttons, and lists.

[0272] "Communication terminal" refers to an electronic device that has display function, input function and data transmission and reception function with a server through a communication network, including but not limited to computing terminal, mobile terminal, tablet terminal, etc.

[0273] "Storage device" refers to an information storage component used to store data such as prompts, version information, and editing history, including local storage and remote storage systems connected via a network.

[0274] "Version information" refers to the identification data corresponding to the status of each prompt statement, used to distinguish prompt statements at different times or different modification stages, including version identifier, timestamp, etc.

[0275] "Feedback information" refers to structured or semi-structured information generated by a generative AI model or a generative AI model used for evaluation based on a customized prompt statement, including evaluation of the advantages of the prompt statement, suggestions for improvement, and revision plans.

[0276] "Revision scheme" refers to one or more rewritten prompt statement candidates generated based on the customized prompt statement and model analysis results, which are used as alternative text for further editing by the user or for direct use.

[0277] "Candidate prompts" refer to prompts that users can choose from among multiple revision options extracted from feedback information. Each candidate prompt can be selected by the user and used as a new editing basis.

[0278] "Edit history" refers to the record information of each version of the prompt statement generated during the process of the user making multiple modifications to the prompt statement, including the content of each version of the prompt statement and its associated time information or operation information.

[0279] "Historical information" refers to a set of recorded data used to characterize the status of prompt statements at different times or different editing stages, including editing history, version information, and operation logs related to prompt statements.

[0280] "External generative AI model" refers to a generative AI model that is not directly integrated into the system but is invoked through a network interface or external service, and is used to generate target content based on the final prompt statement.

[0281] "External information processing device" refers to a computing device that communicates with this system via a network and performs further data processing or application on the final prompt statement, including server systems, business systems or other data processing platforms.

[0282] "Natural Language Processing Parsing" refers to the process by which a server processes text information obtained from a user, performing word segmentation, part-of-speech tagging, keyword extraction, and semantic understanding, in order to obtain a structured or semi-structured representation that can be used by generative artificial intelligence models.

[0283] "Multi-stage interaction" refers to the interactive process between users and servers around prompt statements, which includes multiple recurring cycles such as automatic generation, user editing, model feedback, and further editing.

[0284] The embodiments of this invention will focus on the server, terminal, and user as the main components, and will describe the hardware configuration, software configuration, data structure, and specific data processing and calculation methods of the system. It will also illustrate how this invention can generate and optimize prompt statements within a computer, thereby improving the input construction accuracy, processing efficiency, and data management capabilities of generative artificial intelligence models, in conjunction with several alternative implementation methods.

[0285] I. Overall System Composition In this invention, the server operates as the core data processing device. The server can be an information processing device configured with a multi-core central processing unit (CPU) and a graphics processing unit (GPU), such as a multi-core general-purpose processor, a graphics processor supporting general-purpose parallel computing, and running a general-purpose operating system (such as a Unix-like operating system or other server operating system). The server installs and runs the following software components in its memory: (1) Web application server components (e.g., request processing modules based on HTTP / HTTPS protocols, which can be implemented using a general web framework); (2) Natural language processing library components (e.g., Chinese processing libraries based on word segmentation, part-of-speech tagging and dependency parsing, such as general word segmentation libraries, syntactic analysis libraries, etc.); (3) Generative artificial intelligence model components (e.g., pre-trained language models based on the Transformer architecture, including large-scale neural networks with encoder-decoder structures or decoder-only structures). (4) Data storage components (such as relational database systems, key-value database systems, and file storage systems) are used to store structured or semi-structured data such as prompt statements, version information, and editing history; (5) API interface management component, used to manage message format and calling protocol between server and terminal.

[0286] As a device for user interaction, a terminal can be a general-purpose computing device or mobile device equipped with a display and input devices, such as a mobile terminal with a screen and touch input, or a computing terminal with a keyboard and monitor. Applications or browser applications running locally on the terminal include: (1) User interface rendering module, used to display prompts, feedback information and candidate versions; (2) Input acquisition module, used to acquire text input and selection operations from the user; (3) Local storage module, used to cache the prompt statement draft and historical version identifier in the terminal's non-volatile storage medium; (4) Network communication module, used to exchange data with the server through a secure communication protocol.

[0287] Users interact through the terminal, including inputting creative requirements, editing prompts, selecting candidate prompts, and finally determining the prompts to be used by the external generative artificial intelligence model.

[0288] II. Server-side program structure and data processing 1. Server module division The server logically includes the following functional modules, each of which can be implemented through multiple software processes or threads running on the same processor, or deployed in a distributed manner on different physical servers: (1) Request receiving module: used to parse request messages from the terminal and convert the text fields, task type and user identifier in the request body into internal data structures.

[0289] (2) Natural Language Parsing Module: Used to segment, extract keywords and generate semantic representations of natural language descriptions input by users.

[0290] (3) Prompt statement generation module: used to call the generative artificial intelligence model to generate an initial scheme of prompt statements based on structured semantic information.

[0291] (4) Prompt Statement Evaluation and Feedback Generation Module: This module is used to input user-customized prompt statements into the generative artificial intelligence model or a separately deployed evaluation generative artificial intelligence model to generate feedback information including advantages, improvement points, and revision schemes.

[0292] (5) Version Management and History Module: This module is used to associate and store the prompts for each version with the version number, timestamp, and related metadata, and supports the retrieval and rollback of historical versions.

[0293] (6) External model docking module: After the user confirms the final prompt statement, the prompt statement is forwarded to the external generative artificial intelligence model or external information processing device in a predefined format.

[0294] 2. Specific data structures and operations within the server A server maintains multiple logical data tables or collections in its storage device. For example: (1) Prompt Statement Table: Includes fields such as "Prompt Statement Number", "User Number", "Current Content", "Language Type", and "Associated Task Type".

[0295] (2) Version table: includes fields such as "Version number", "Prompt statement number", "Version sequence number", "Time stamp", "Content snapshot", and "Source type (initial generation / user editing / model revision)".

[0296] (3) Feedback Form: Includes fields such as "Feedback Number", "Version Number", "Advantages Text", "Improvement Points Text", and "Revision Plan List".

[0297] The server uses word segmentation algorithms to process user input descriptions in its natural language parsing module. The server can employ a word segmenter combining dictionary matching and statistical language models, or a word segmentation model based on bidirectional long short-term memory networks and conditional random fields. During semantic representation, the server maps words to vector representations and performs multi-layer transformations to obtain the contextual representation of the text for use by the subsequent prompt generation module.

[0298] The server uses a Transformer-based neural network when invoking the generative AI model. This network consists of multiple self-attention sublayers and feedforward sublayers. Each self-attention sublayer calculates attention weights through linear transformations of the query, key, and value. The feedforward sublayer transforms the information through two fully connected layers combined with a non-linear activation function. During the inference phase, the server encodes the input text into a series of discrete tokens, which are then converted into continuous vectors through an embedding layer. At each time step, the server calculates the probability distribution of the next token based on the context vector from the previous time step and generates continuous text output using sampling strategies (such as temperature sampling, Top-k, or Top-p sampling).

[0299] During the evaluation and feedback generation phase, the server concatenates user-customized prompts and evaluation instructions into a new input sequence, which is then fed into the generative AI model. The server explicitly includes instructions such as "Please point out the advantages" and "Please provide at least one revised version" in the input, ensuring that the model's output has multiple parsable parts. After receiving the model output, the server parses it using keywords or delimiters, separating the advantages, improvements, and revisions into independent fields and writing them into the feedback table.

[0300] The server performs numerical and vector operations during these processes, such as matrix multiplication, weighted vector summation, and nonlinear transformations. By performing these matrix operations in batches on the GPU, the server can significantly improve the parallelism of prompt generation and evaluation calculations, thereby increasing processing speed and system throughput.

[0301] III. Program Structure and Display Method on the Terminal Side In this invention, the terminal is responsible for directly interacting with the user and presenting the prompts and feedback information generated by the server in an easy-to-understand and editable manner.

[0302] The terminal divides the user interface into multiple areas, for example: (1) Input Request Area: This area allows users to input the content topic or request description they want to generate, such as "I want to write a medieval fantasy novel about knights fighting".

[0303] (2) Prompt statement editing area: used to display the initial scheme of prompt statements generated by the server and allow users to edit the text via keyboard or touch.

[0304] (3) Feedback display area: used to display the advantages, improvements and revisions returned by the server, and to display candidate prompts in the form of a list or card.

[0305] (4) Historical Versions Area: This area displays the version number and time information of the current prompt statement, making it easy for users to switch and compare between different versions.

[0306] The terminal maintains a lightweight data structure in local non-volatile storage to cache the current user's prompts and recent version identifiers, allowing it to resume editing even after a temporary network outage or application closure. This local caching reduces frequent read requests to the server, thereby decreasing communication load and improving overall system response speed.

[0307] IV. Examples of specific user operations in the system Users input their creative requirements into the terminal, for example: "I want to write a fantasy story about brave battles." After receiving the input, the server uses natural language parsing to construct an initial prompt statement based on key concepts such as "brave battle," "fantasy story," "medieval," and "knight." An example prompt statement generated by the server is as follows: "Describes a medieval knight fighting bravely on a smoke-filled battlefield." The terminal displays this prompt in the editing area, which the user can modify to: "Describes a medieval knight in tattered silver armor, charging into enemy lines on a smoke-filled battlefield at dusk, shouting as he fights bravely." After receiving the user-customized prompt, the server stores the prompt as a version and inputs it into the generative AI model used for evaluation. The server adds the following instruction to the text input to the model: Please evaluate the following prompts, pointing out their advantages and areas for improvement, and provide at least one rewritten version.

[0308] Prompt: Describes a medieval knight in tattered silver armor, charging bravely into enemy lines on a smoke-filled battlefield at dusk, longsword raised high, shouting as he fights. The server outputs structured content similar to the following: Advantages: The time and environment are clearly described, and the actions are vividly portrayed.

[0309] Improvements could include adding sensory details and depicting the characters' inner thoughts.

[0310] Revision proposal: "On the smoke-filled battlefield at dusk, a medieval knight in tattered silver armor charged into the enemy ranks, sword raised high. Firelight illuminated his sweat-drenched face, and all around him were shouts that tore through the air and the clang of weapons. He shouted loudly while suppressing his inner fear, fighting bravely to protect the fragile kingdom behind him." The terminal displays the revised solution as candidate prompts, which users can select with a single click and fine-tune further, such as adding descriptions like "cinematic" or "high detail." The end user receives prompts that can be directly used with external text generation models, for example: "Write a medieval fantasy novel where the protagonist is a young knight wearing slightly worn silver armor. On a battlefield where fire and smoke intertwine, and the sounds of clashing metal and screams rise and fall, he grips his longsword tightly, suppressing the trembling in his body, and fights bravely to protect the kingdom on the verge of destruction and his lover who awaits him in the distance." Users can also use a similar approach in image generation scenarios to construct the following prompt statement: "A medieval knight in tattered silver armor, wielding a longsword, charges across a smoke-filled battlefield at dusk. Fantasy style, warm orange lighting, high detail, cinematic feel." The server can directly forward this prompt to an external image generation model to generate the corresponding image result.

[0311] V. Improvements and Effects of Computer Technology In this invention, the server not only executes the traditional simple process of "receiving text - returning text," but also improves computer technology itself through the following mechanism: (1) By introducing version management of prompt statements and structured feedback generation, the server forms a traceable evolution trajectory of prompt statements in the database, which enables subsequent model optimization and user preference modeling to use these historical data for statistical analysis, thereby improving the stability and repeatability of the input construction of generative artificial intelligence models.

[0312] (2) During the natural language parsing and prompt generation process, the server adopts multi-stage feature extraction and vector representation, which enables prompts with similar needs to be clustered in the vector space. This is beneficial for caching and reuse on the server side, reducing redundant calculations and lowering the overall inference load.

[0313] (3) Through the structured analysis of feedback information, the server internally forms clear fields such as "advantages", "improvements" and "revision schemes". These fields can be indexed, retrieved and combined separately, thereby supporting subsequent automatic recommendation and semi-automatic construction algorithms and improving human-machine collaboration efficiency.

[0314] (4) The server can control the temperature parameters and Top-k or Top-p sampling range of the generative AI model to adopt different exploration-exploitation strategies at different stages: more emphasis is placed on diversity in the initial generation stage and more emphasis is placed on stability in the evaluation and revision stage. This dynamic control enables the system to reduce invalid generation while ensuring generation quality, thereby improving the computational efficiency of the entire system.

[0315] (5) During the learning phase, the server can perform supervised learning on a large number of archived "initial prompt statement – ​​user-edited version – final prompt statement" triples, using cross-entropy loss function or sequence-to-sequence loss function to adjust model parameters, thereby making the model closer to the editing mode of real users. During backpropagation, the server updates the weight matrix of each layer of the network according to the loss function, and uses optimization algorithms (such as first-order gradient-based optimization algorithms) to achieve parameter iteration. Through this training method, the model can internally learn a structured expression that is different from human editing habits but more suitable for unified computer processing, thereby reducing the editing burden on users.

[0316] The above-mentioned processing makes the system of the present invention no longer a simple human-operated automated tool, but rather a system that constructs a dedicated data structure and multi-stage algorithm flow within the computer to address the problem of prompt statement design. This systematically optimizes the input construction of the generative artificial intelligence model, thereby achieving technical improvements in terms of generation quality, response speed, and resource utilization.

[0317] VI. Alternative Implementation Methods and Extensions The server can employ different generative artificial intelligence model structures in different implementations. For example: (1) The server can use an autoregressive language model with a decoder-only structure to generate prompts word by word; (2) The server can use a sequence-to-sequence model with an encoder-decoder structure to encode the long descriptions input by the user into context vectors and then decode them into compact prompt statements; (3) The server can adopt a multi-task training model to learn two tasks, "prompt statement generation" and "feedback generation", in the same network at the same time, and improve parameter utilization by sharing the encoding layer and separating the decoding head.

[0318] The terminal may also have different interface layouts and local storage strategies in different implementations, for example: (1) The terminal can provide only local version management and editing functions in offline mode, and then synchronize to the server in batches when the network is available; (2) The terminal can dynamically adjust the layout of the prompt editing area and the feedback display area according to the screen size, so that users can still effectively browse and edit on small screen devices; (3) With user authorization, the terminal can sort the frequently used prompts locally and prioritize the display of frequently used templates to further reduce the number of server calls.

[0319] Through the above-mentioned various implementation forms and alternative methods, this invention takes into account both system flexibility and scalability while ensuring functional integrity, enabling efficient collaboration among the server, terminal and user in the process of constructing prompt statements in the generative artificial intelligence model, thereby bringing technical effects such as improved processing speed, improved accuracy of prompt statement construction and enhanced data management capabilities at the computer level.

[0320] use Figure 13 The processing procedure is explained.

[0321] Step 1: Users input their requirements on the terminal and send them to the server.

[0322] Users enter their natural language requirements in the terminal's text input area, such as "I want to write a fantasy story about brave battles" or "I want to generate an illustration of a medieval knight in battle."

[0323] The terminal takes the natural language text input by the user as input data, calls the local input acquisition module, and combines the text with the task type (such as "text generation" or "image generation description") to form structured request data. The terminal encodes the data (e.g., UTF-8 encoding) and encapsulates it into a request message, which is then sent to the server as an HTTPS request via the network communication module.

[0324] The input for this step is the natural language text and task type entered by the user, and the output is a structured request message sent to the server; during this process, the terminal performs basic formatting and encoding on the text so that the server can reliably parse it.

[0325] Step 2: The server parses the request and performs natural language preprocessing.

[0326] The server receives request messages from the terminal in the request receiving module, taking the text fields in the message body and the task type as input. The server uses a parser to unpack the message and parse the JSON, extracting fields such as "requirement description text", "task type", and "user identifier".

[0327] The server inputs the "requirement description text" into the natural language parsing module, calls a word segmentation algorithm to divide the continuous text into a sequence of words, uses a part-of-speech tagging algorithm to assign a part-of-speech tag to each word, and uses a keyword extraction algorithm (e.g., based on TF-IDF or attention weight statistics) to select several high-weight keywords. The server then converts the word sequence into a vector sequence through an embedding layer, and generates an overall semantic representation vector through a multi-layer neural network or a pre-trained encoder.

[0328] The input to this step is structured request data containing natural language requirements, and the output is an intermediate data structure containing a keyword list, a word sequence, and a semantic vector. The server processes the data through word segmentation, part-of-speech tagging, keyword extraction, and vectorization to obtain semantic features suitable for subsequent generation and processing.

[0329] Step 3: The server generates initial instructions by constructing prompt statements and then invokes the generative artificial intelligence model.

[0330] The server takes the keyword list, requirement description text, and task type obtained in step 2 as input and calls the prompt statement generation module to construct the model instruction text. For example, the server generates the following template text: "Please generate a prompt statement suitable as input for a generative artificial intelligence model based on the following themes and keywords."

[0331] Theme: Fight Bravely Keywords: Middle Ages, knights, battlefield. The server takes the template text as input, converts it into a sequence of tag IDs through word segmentation and tokenization, and then inputs it into a generative AI model based on the Transformer architecture. The model performs matrix multiplication, self-attention calculation, and feedforward network operations on the GPU, progressively predicting the probability distribution of the next tag based on the current context vector. The server samples according to the set generation parameters (such as maximum length, temperature, Top-k, or Top-p) until the final tag is generated.

[0332] The input for this step is structured semantic features and template text, and the output is one or more candidate prompt statement strings. The server uses neural network inference and probability sampling operations to convert the abstract semantic features into an initial scheme of natural language prompt statements.

[0333] Step 4: The server selects the initial prompt statement and sends it to the terminal.

[0334] The server takes one or more candidate prompts output by the generative AI model as input and sorts them using rules such as length, coherence score, or internal confidence. The server can select the prompt with the highest score as the "initial prompt scheme," for example: "Describes a medieval knight fighting bravely on a smoke-filled battlefield." The server assembles the selected prompt statement with metadata (such as model name and generation time) into a response message, and sends it to the terminal as an HTTPS response through the request-response module.

[0335] The input to this step is the set of candidate prompt statements generated by the model, and the output is a response message containing the selected initial prompt statement. The server reduces redundant candidates through sorting and filtering operations, and only outputs the initial solutions with higher quality.

[0336] Step 5: The terminal displays an initial prompt and allows users to edit it.

[0337] The terminal receives the server's response message in the network communication module, takes the initial prompt statement field as input, and passes it to the user interface rendering module. The terminal displays the prompt statement on the screen in the editable text area and shows explanatory information on the interface to guide the user in making modifications.

[0338] Users can edit the prompts on the terminal using keyboard or touch input, for example, by... "Describes a medieval knight fighting bravely on a smoke-filled battlefield." Revised to: "Describes a medieval knight in tattered silver armor, charging into enemy lines on a smoke-filled battlefield at dusk, shouting as he fights bravely." The terminal takes the text edited by the user as input each time, updates the current prompt statement content in the local memory, and writes the content to the local storage module as a draft when necessary.

[0339] The input for this step is the initial prompt returned by the server and the user's input, and the output is the currently customized prompt text. The terminal updates the prompt locally through interface display and text editing buffer operations.

[0340] Step 6: The terminal submits a customized prompt message to the server and requests feedback.

[0341] After the user clicks the "Get Feedback" button or a similar button, the terminal takes the prompt text in the current editing area as input, and packages it together with metadata such as user identifier and task type into a request message. The terminal performs necessary escaping and encoding on this text, and sends it as the request body to the server's feedback generation interface via HTTPS.

[0342] The input for this step is the customized prompt text and related metadata, and the output is a feedback request message sent to the server; the terminal performs data encapsulation and network transmission preparation at this stage.

[0343] Step 7: The server stores versions of custom prompt statements and constructs evaluation instructions.

[0344] After receiving the terminal's request through the feedback interface, the server takes the customized prompt text as input, calls the version management module to generate a new version number for the prompt, and writes information such as "prompt text number", "version number", "content snapshot", "timestamp", and "source type (user edit)" into the database.

[0345] The server then constructs evaluation instruction text, inserting the customized prompt statement into a template that requires structured feedback, for example: "You are a prompt statement optimization assistant."

[0346] Task: Evaluate and improve the following prompts to make them more suitable as input for generative artificial intelligence models.

[0347] Require: 1. Point out the advantages of the current prompt statement; 2. Point out areas for improvement; 3. Provide at least one rewritten version of the prompt statement.

[0348] Current prompt: Describes a medieval knight in tattered silver armor, charging into enemy lines on a smoke-filled battlefield at dusk, sword in hand, shouting as he fights bravely. The input for this step is the customized prompt statement and version metadata, and the output is the version record written to the database and the constructed evaluation instruction text. The server ensures version traceability through write operations and constructs a unified format evaluation input through text concatenation operations.

[0349] Step 8: The server invokes a generative artificial intelligence model to generate feedback information.

[0350] The server takes the evaluation instruction text constructed in step 7 as input, and after tokenization and vectorization, feeds it into the generative artificial intelligence model for evaluation. The model performs forward propagation in a multi-layer Transformer network, calculates the probability distribution of the output label at each time step, and the server generates continuous text containing "advantages", "improvements", and "revision schemes" according to the sampling strategy.

[0351] After receiving the model output, the server takes the complete output string as input and calls the text parsing module to segment it according to predefined delimiters or keywords (such as "Advantages:", "Improvements:", "Revision Scheme:"). Different parts are mapped to corresponding fields to form a structured feedback data object. The server then stores this feedback object along with its corresponding version number in the feedback table.

[0352] The input for this step is the evaluation instruction text, and the output is structured feedback information (including text of advantages, text of improvements, and a list of revision schemes). The server uses data operations such as neural network reasoning and rule parsing to convert the unstructured generated text into directly usable structured feedback data.

[0353] Step 9: The server returns feedback information to the terminal.

[0354] The server reads the newly generated feedback record from the feedback table, takes the "Advantages," "Improvements," and "Revision Suggestions" fields as input, and assembles them into a response message body. The server can include multiple revision suggestions in the message as candidate prompts.

[0355] The server sends the feedback message to the terminal via HTTPS through the response module.

[0356] The input for this step is the feedback field stored in the database, and the output is the feedback response message sent to the terminal; the server performs data reading and message encapsulation processing.

[0357] Step 10: The terminal displays feedback and allows users to select revision options.

[0358] The terminal receives feedback messages from the server, taking "Advantages," "Improvements," and "List of Revisions" as input, and passes them to the interface rendering module. The terminal displays the following on the screen: – Strengths (e.g., "clear description of time and environment, and vivid description of actions"); – List of improvements (e.g., “We could add sensory details and character inner thoughts”); – Multiple revision options, each displayed as an independent candidate suggestion statement.

[0359] Users can click on a revision scheme on the terminal. The terminal uses the clicked text as input, overwrites the current editing area, and marks the revision scheme as a "from model" version. Users can also continue to manually modify the text content based on this.

[0360] The input for this step is feedback information and the user's selection, and the output is the updated content of the editing area and the new current prompt text. The terminal transforms the structured feedback into a text object that the user can directly manipulate through interface interaction and content replacement operations.

[0361] Step 11: Users perform multiple rounds of iterative editing and finally determine the prompt message.

[0362] Users can continue to edit the prompts on the terminal based on feedback, for example, adding requirements such as "cinematic feel" and "high detail" to the revised version, resulting in more complete prompts, such as: "Write a medieval fantasy novel where the protagonist is a young knight wearing slightly worn silver armor. On a battlefield where fire and smoke intertwine, and the sounds of clashing metal and screams rise and fall, he grips his longsword tightly, suppressing the trembling in his body, and fights bravely to protect the kingdom on the verge of destruction and his lover who awaits him in the distance." Each time the user confirms "Get new feedback", the terminal sends the current text as input to the server again. The server repeats the processing from steps 7 to 10, thus forming multiple rounds of iteration.

[0363] When the user believes that the current prompt statement meets their needs, the user performs the "Confirm Final Prompt Statement" operation on the terminal, and the terminal submits the final text as input to the server or copies it to the clipboard.

[0364] The input for this step is the text of the prompt statement after each round of editing by the user, and the output is the final prompt statement confirmed by the user. Through multiple rounds of feedback-editing loops, the system generates high-quality prompt statements that have undergone multiple machine evaluations and human revisions.

[0365] Step 12: The server or terminal will apply the final prompt statement to the external generative artificial intelligence model.

[0366] When the user selects "Generate content using this prompt," the terminal takes the final prompt as input and can perform one of two operations: – Copy the text to the terminal clipboard, and have the user manually paste it into the external generative AI model interface; – The final prompt statement is encapsulated into a request message through the network communication module and sent to the external model interface of the server.

[0367] After receiving the final prompt, the server uses that text as input, constructs a request message according to the interface specification of the external generative AI model, and sends it to the external model service. The external model performs inference and returns the generated result (such as a reference address to text fragments or image data), which the server then forwards to the terminal.

[0368] The input for this step is the final prompt text, and the output is the target content or its access information generated based on the prompt. The server and terminal, through the interface with the external model, realize the complete technical process from prompt construction to content generation.

[0369] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0370] The following technical problems exist in existing content generation and interaction systems: First, servers typically only call generative AI models to generate fixed prompts or content based on the user's one-time text input, lacking fine-grained tracking and structured utilization of the user's subsequent editing behavior. This results in the inability to iteratively optimize the personality settings, making it difficult for the generated results to consistently match the user's true creative intent.

[0371] Second, servers generally treat user input as static data and rarely consider the user's emotional state during use as a real-time input signal in the calculation process. This makes it impossible for the system to adaptively adjust the tone and content focus of the prompts during the model inference stage, thus making it difficult to alleviate the user's stress, frustration and other negative experiences during the interaction process in a timely manner.

[0372] Third, traditional recommendation systems mostly rely on click records or rough interest tags for content retrieval, without jointly modeling the semantic information contained in the personality prompts and the user's current emotional state. This results in the recommendation algorithm being unable to provide highly relevant personalized content for creative needs or emotional states in specific contexts, thus reducing the actual auxiliary value of the recommendation results for the creative process.

[0373] Fourth, existing systems typically separate "character personality setting" from "content recommendation logic," lacking a unified computational framework that allows generative AI models to both generate or adjust personality prompts and generate internal search condition prompts, thereby driving downstream content retrieval and recommendation. This fragmented architecture increases data conversion costs between modules in engineering implementation and limits the overall system's scalability and maintainability.

[0374] Fifth, many creative aids only provide static examples or templates, failing to offer "personality generation instructions" based on generative artificial intelligence models as a high-level creative reference. Users lack structured guidance information when conceiving new virtual objects, often requiring repeated trial-and-error input, which wastes computing resources and reduces the efficiency of human-computer interaction.

[0375] Therefore, an improved computer implementation is needed to enable the server to: (1) integrate user input, editing behavior and emotional state into a unified information processing framework; (2) dynamically generate and adjust personality prompt statements; (3) automatically generate internal prompt statements for retrieval and recommendation based on generative artificial intelligence models; and (4) provide structured generation instructions for users to create new personality prompt statements, thereby substantially improving the quality of content generation, recommendation accuracy and interactive experience, and achieving an overall improvement in the performance of computer technology.

[0376] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.

[0377] In this invention, the server includes an information processing unit configured to: utilize a generative artificial intelligence model to perform natural language processing parsing on user input information obtained from a communication terminal, automatically generating personality prompt statements to instruct virtual object behavior or reactions; output the personality prompt statements to the display area of ​​the communication terminal, and obtain user editing or appending operations on the personality prompt statements, updating the personality prompt statements according to the operations, so that the personality prompt statements can be iteratively personalized; perform emotion estimation processing based on image or sound information sent by the communication terminal, obtain the user's emotional state in chronological order, and use the emotional state and personality prompt statements together as input to execute the generative artificial intelligence model again, thereby generating adjusted personality prompt statements that dynamically adjust the content or tone of expression according to the emotional state; further, the information processing unit is configured to: utilize a generative artificial intelligence model to perform natural language processing parsing on user input information obtained from a communication terminal, automatically generating ... The information processing unit is also configured to generate content prompts, including virtual object dialogue statements, behavioral description statements, or interactive story branch information, based on the adjusted personality prompts and send them to the communication terminal to provide interactive information presentation or creative assistance; furthermore, it uses the user's emotional state and the words contained in the personality prompts as features to search the content storage area, extracting information content that matches the emotional state to generate recommended information; simultaneously, the information processing unit also generates internal search condition prompts related to the user's emotional state and personality prompts through a generative artificial intelligence model, extracts candidate content from the content storage area based on these search condition prompts, and when a new virtual object creation request is received, generates a personality generation instruction statement containing the virtual object's role, internal contradictions, and story development strategy as creative reference information and presents it to the communication terminal. This allows for the formation of a unified information processing flow centered on a generative artificial intelligence model within the server. This process organically combines user input, editing behavior, emotional state, and content retrieval, enabling adaptive generation and dynamic adjustment of personality prompts. It improves the relevance and diversity of content generation and recommendation, reduces redundant data conversion between multiple modules, and enhances the system's response efficiency and computing resource utilization in interactive creation scenarios. Ultimately, this improves computer-based content generation and recommendation technology as a whole.

[0378] "System" refers to a whole consisting of one or more information processing devices, communication terminals and software modules running on them, which is a computer implementation device used to perform processing such as generating and adjusting personality prompt statements and recommending content.

[0379] "Information processing device" refers to an electronic device with a processor, memory and communication interface, which may be a server, cloud computing node or other computer hardware capable of executing program instructions.

[0380] "Information processing unit" refers to a set of program modules or functional modules that run on an information processing device to perform specific data processing logic, including but not limited to processing functions such as parsing user input, calling generative artificial intelligence models, performing sentiment estimation, and performing content retrieval.

[0381] "Communication terminal" refers to user-side equipment that can communicate data with information processing devices via wired or wireless networks, including but not limited to mobile terminals, fixed terminals, display devices, or wearable devices.

[0382] "User input information" refers to text data, control instructions, or other symbolic data provided by the user through the input interface of the communication terminal to describe virtual objects, scenes, or creative intentions.

[0383] "Generative artificial intelligence models" refer to artificial intelligence models that are trained using machine learning algorithms and can generate text, vectors, or other data outputs based on input data, especially text generation models based on natural language processing technology.

[0384] "Natural Language Processing (NLP) parsing" refers to the technical process of processing user input information through word segmentation, part-of-speech tagging, syntactic analysis, and semantic recognition in order to extract keywords, semantic structures, or contextual information.

[0385] "Virtual objects" refer to characters, groups of characters, or abstract personality entities set in digital content or interactive scenarios for use in system behavior simulation, dialogue generation, or story interpretation.

[0386] "Personality cue statements" refer to textual descriptions used to indicate the behavior, reaction patterns, language style, or value orientation of virtual objects, serving as conditional inputs for generative artificial intelligence models in subsequent content generation.

[0387] "Automatic generation" refers to the process by which an information processing unit automatically generates output data based on input data, without requiring manual writing of each line. This is achieved by executing predetermined program logic and calling generative artificial intelligence models.

[0388] "Editing or adding operations" refers to the changes made by users on the communication terminal, such as modifying, deleting, adding, or rewriting the displayed personality prompts.

[0389] "Personalized customization" refers to the process and result of adjusting the text content of personality prompts based on the user's editing operations and preferences, so that the personality prompts reflect the specific user's creative needs and aesthetic preferences.

[0390] "Image information" refers to image data or image sequence data collected and transmitted by the imaging device of a communication terminal to reflect the user's facial expressions or posture.

[0391] "Audio information" refers to audio data collected and transmitted by the microphone of a communication terminal to reflect the user's voice characteristics, tone, or emotional state.

[0392] "Emotion estimation processing" refers to the process of inferring a user's current emotional state by extracting features and recognizing patterns from image information, sound information, or other physiological signals.

[0393] "Emotional state" refers to the classification result or continuous quantity that represents a user's psychological or emotional state within a specific time period, obtained through emotion estimation processing, including but not limited to joy, sadness, stress, frustration, excitement, etc.

[0394] "Adjusted personality prompt statements" refer to personality prompt statements that are dynamically modified based on the original personality prompt statements, according to the emotional state obtained by the information processing unit, and through a generative artificial intelligence model to modify the content, tone, or key information of the statements.

[0395] "Statement content" refers to the specific words and semantic structure used in personality prompt statements to describe the nature, behavior, and reaction patterns of virtual objects.

[0396] "Expressive tone" refers to the stylistic tendency of personality cues in language statements, including but not limited to calm, gentle, encouraging, humorous, and serious tone characteristics.

[0397] "Content prompt statements" refer to text outputs generated based on the adjusted personality prompt statements, used directly as virtual objects to speak, describe behavior, or describe interactive story branches.

[0398] "Speech lines" refer to the dialogue text that virtual objects are set to output in a dialogue scenario, used for language interaction with users or other virtual objects.

[0399] "Behavioral description statements" are texts used to describe the actions, decisions, or mental activities that a virtual object will perform or is performing in a specific context.

[0400] "Interactive story branching information" refers to structured text descriptions used in interactive narrative scenarios to indicate different story directions, choice outcomes, or subsequent plot nodes.

[0401] "Information presentation" refers to the process of displaying content prompts, recommendations, or creative references to users in a visual or auditory form through the display device or audio output device of a communication terminal.

[0402] "Creative assistance" refers to the function of helping users to conceive, modify, or expand virtual objects and their related story content by providing personality prompts, content prompts, or personality generation instructions.

[0403] "Words" refer to the linguistic units that make up the text in personality prompt statements, including single words, phrases, short phrases, or other text fragments with independent semantic meaning.

[0404] "Features" refer to the numerical representations extracted from personality prompts and emotional states during content retrieval and recommendation, used for modeling and calculation, including but not limited to keyword vectors, emotion tag encodings, and statistical features.

[0405] "Content storage area" refers to the data storage space that stores various types of information content that can be retrieved. It can be a database, file storage system or other data management system.

[0406] "Information content" refers to text, audio, video, image, or mixed media data stored in the content storage area that can be recommended to users.

[0407] "Recommendation information generation unit" refers to a functional module or program component implemented in an information processing device, which is used to retrieve content storage areas based on feature values ​​and generate recommendation results.

[0408] "Internal search condition prompts" refer to text descriptions generated by generative artificial intelligence models and used as input content storage areas for search conditions. These descriptions take into account the semantic information of the user's emotional state and personality prompts.

[0409] "Candidate content" refers to a collection of multiple pieces of information that have not yet undergone final filtering after the content storage area has been searched based on internal search criteria.

[0410] "New Virtual Object Creation Request" refers to an instruction or input data sent by a user through a communication terminal, requesting the system to generate initial settings or creative references for a virtual object for which no personality characteristics have yet been set.

[0411] "Personality generation instruction statements" refer to high-level text descriptions generated by generative artificial intelligence models to guide users in constructing the personality settings of virtual objects. They typically include elements such as character positioning, internal conflicts, and story development strategies.

[0412] "Creative Reference Information" refers to personality generation instructions, example scenarios, or other structured suggestions that users can refer to and expand upon during the creative process.

[0413] In the following embodiments, the server, terminal, and user each assume different functional roles. This invention is not limited to the specific hardware and software implementations described below, but for ease of understanding, the following description is based on typical configurations.

[0414] I. Overall System Composition The server includes a processor, main memory, persistent storage, and a network interface. The server can employ a multi-core general-purpose processor, such as a multi-core central processing unit based on the x86 architecture, and may optionally be equipped with a graphics processing unit to accelerate deep learning inference. The server runs multiple software components on an operating system (such as a Linux kernel-based server operating system), including: a web service module, a natural language processing module, a generative artificial intelligence model inference module, a sentiment estimation module, a content retrieval module, and a data management module.

[0415] The terminal includes at least one processor, memory, display device, input device, image acquisition device, and sound acquisition device. The terminal can be a smartphone, tablet, head-mounted display device, or desktop computing device. The client program running on the terminal can implement a graphical user interface based on a native application framework or browser execution environment, used to display personality prompts, content prompts, and creative reference information to the user, and to collect user-input text, editing operations, image information, and sound information.

[0416] Users interact with the server through a terminal. Users input their needs for the virtual object into the terminal's text input controls, edit the personality prompts generated by the server, read the virtual object's dialogue in the dialogue interface, and view personality generation instructions in the creation interface. Users also authorize the camera and microphone to capture facial expressions and voice recordings after authorization, allowing the server to estimate their emotional state.

[0417] II. Generative Artificial Intelligence Models and Natural Language Processing Implementation Forms The server deploys a generative AI model in its inference module. The server can employ a sequence-to-sequence generation model based on a multi-layer self-attention structure, comprising a word embedding layer, a multi-layer encoder-decoder stack, and an output layer. During training, the server compares the model's output labeled probability distribution with the target sequence using a loss function (e.g., cross-entropy loss). In each training batch, gradients are calculated using backpropagation, and the model weights are updated using parameter update algorithms (e.g., optimization algorithms based on adaptive learning rates). During training, the server can employ data augmentation techniques, such as synonym replacement and sentence rearrangement of the original corpus, to improve the robustness of the model in generating prompts across different styles and diverse contexts.

[0418] During runtime, the server encodes user input and related context sequences into discrete token sequences, maps them to continuous vectors via a word embedding matrix, and then feeds them into a multi-head self-attention encoding layer. In attention computation, the server generates query, key, and value vectors for each token, calculates attention weights through scaling dot products, and linearly combines the value vectors with these weights, thereby capturing long-distance dependencies and relationships between keywords within the model. This structure offers significant advantages in context modeling capabilities and generation quality compared to traditional fixed-window-based n-gram models.

[0419] The server uses word segmentation and dependency parsing tools in its natural language processing module to preprocess the user input text, extracting entity names, adjective phrases, and behavior-related predicate phrases. The server converts these semantic units into structured features, which are then used as prefixes for the generative AI model. This better constrains the scope of the generated prompts and reduces irrelevant content. Because this preprocessing employs dependency-structure-oriented feature extraction rather than simple word frequency statistics, it improves the accuracy of generated instructions even with complex, long sentence inputs.

[0420] III. Generation and Updating of Personality Tips After receiving user input from the terminal, the server first performs natural language parsing, converting the user description into an internal semantic representation. Based on this, the server constructs an initial prompt, such as: "Generate personality prompts for a character based on the following user requirements."

[0421] User requirement: 'A brave female swordsman fights in a futuristic city.' The personality cues should describe: speaking style, behavioral characteristics, values, and typical scenarios. The server uses this prompt as the input sequence for a generative artificial intelligence model. During the decoding phase, the server generates an output sequence word by word or character by character using an autoregressive approach until a termination condition is met. The generated result is a personality prompt statement, for example: "This character is a female swordsman living in a high-tech futuristic city. She is brave and decisive, speaks concisely and directly, and dislikes beating around the bush. In battle, she always stands in front of her teammates, protecting them with her actions. She is used to encouraging her companions with a short and powerful sentence at crucial moments, such as: 'Don't be afraid, I'm here.' She hates injustice and has zero tolerance for bullying the weak." After receiving the personality prompt, the terminal displays it to the user in an editable text component. The user can directly modify some wording or add behavioral details, such as adding: "She would always crack a joke after a battle to ease the tension, such as making fun of her disheveled appearance." When a user submits modifications, the terminal sends the edited, complete personality prompt statement back to the server. The server, in its data management module, stores the before and after prompt statements in a structured format in storage, such as recording the prompt content, version number, and associated emotional state tags in a relationship table. When generating subsequent content, the server prioritizes using the latest version of the personality prompt statement, ensuring that the model's inferences are based on settings actually accepted by the user.

[0422] In this way, the server no longer treats user input as a one-time instruction, but models it together with the user's editing behavior as a "personality configuration state". This structured multi-version storage and associated usage method allows the server to continuously inherit and adjust the role settings between multiple inference calls, thereby achieving a stable generation result that meets user expectations. This reduces the workload of users re-entering or repeatedly fine-tuning, and also reduces the redundant calculation overhead of the server in multiple rounds of interaction.

[0423] IV. Emotion Estimation and Adjustment of Emotion-Driven Cue Statements When acquiring image information, the terminal uses a camera to capture sequences of facial images of the user and encodes the images at fixed time intervals. The terminal can downsample and compress the images locally to reduce data size. Similarly, the terminal uses a microphone to capture segments of the user's speech and extracts basic acoustic features such as short-time energy, fundamental frequency, formants, and speech rate. The terminal packages this feature data and transmits it to the server via a network interface.

[0424] The server deploys a classification structure combining convolutional neural networks (CNNs) and recurrent neural networks (RNNs) in its sentiment estimation module. For example, the server uses CNNs to extract spatial features from facial image features, employing weights pre-trained on a large-scale sentiment-annotated dataset; for speech features, it uses bidirectional recurrent unit networks (BRNNs) to extract time-series features. The server concatenates or weights the visual and acoustic feature vectors in the fusion layer and outputs the sentiment category probability via a fully connected layer. When training this sentiment estimation network, the server uses a multi-label cross-entropy loss function to predict different sentiment categories simultaneously and updates the network parameters using a stochastic gradient descent-like algorithm.

[0425] During operation, the server applies a threshold or weighted average strategy to the probability distribution of emotions output at each time point to calculate the representative emotional state of the current session. For example, the server can perform an exponentially weighted average of the probabilities of emotion labels such as "stress" and "frustration" over a recent time window to obtain a smooth emotion curve and reduce erroneous reactions caused by momentary fluctuations. The server writes this emotional state along with the current personality prompt into an intermediate feature structure, which is then input into the generative artificial intelligence model.

[0426] When the server detects that a user is under stress, the server constructs an update command, for example: "The following is a personality prompt: 'This AI assistant speaks calmly and directly, often using a commanding tone to demand that users complete tasks on time.'" The user's current emotion is 'stressed'. Please rewrite the prompts to be gentler and more reassuring, while maintaining the assistant's efficiency. The server sends the update command to the generative artificial intelligence model, which automatically generates a new version of the personality prompt during the decoding process, for example: "This AI assistant remains calm and efficient, but it will first consider the user's state before gently reminding them of the task assignment. It will use an encouraging tone, such as: 'You've done a great job. Let's break down the remaining tasks into smaller steps and complete them together.'" The server utilizes a combined input of "emotional state + existing personality cues," rather than simply rewriting the cues themselves. This allows the model's attention to simultaneously focus on both semantic content and emotional control signals, enabling targeted tone adjustments. Compared to traditional post-hoc rule replacement methods, this structured emotion-driven mechanism can simultaneously generate content and adjust style within a single inference iteration, improving output stability and computational efficiency.

[0427] V. Content Prompt Generation and Content Recommendation After obtaining the adjusted personality prompt, the server generates content prompts based on the specific application scenario. The server uses the following commands in virtual character dialogue scenarios: "The following is a statement indicating the character's personality."

[0428] 'This character is a female swordsman living in a high-tech futuristic city... She always cracks a joke to lighten the mood after a battle.' Based on the given prompt, generate three sentences suitable for speaking to the audience during a live broadcast, each no more than 30 characters long. The server inputs this instruction into the generative artificial intelligence model. The model uses learned language patterns and character styles to generate short sentences at the output, such as: "Was that sword strike cool? Remember to give me a like." "Don't worry, I'm still standing. Leave the next enemy to me." "Don't be fooled by my disheveled appearance, I'm still a formidable fighter." The server packages these content prompts and sends them to the terminal. The terminal can display them in the live broadcast control interface for the broadcaster to insert at appropriate times, or they can be played automatically through the speech synthesis module. Unlike traditional fixed scripts, the server generates new content prompts based on the latest personality prompts and emotional states in each call, thereby achieving high diversity of output while maintaining consistency in roles and reducing the cost of manual script writing and maintenance.

[0429] In content recommendation, the server jointly encodes the user's emotional state with keywords in the personality prompts into a high-dimensional feature vector. The server uses word embedding algorithms to map keywords to a vector space, then maps emotion tags to independent vectors, and concatenates or fuses them through linear transformations to form a retrieval vector. The server maintains feature representations of various information contents in the content storage area; for example, it pre-calculates emotion tendency vectors and topic vectors for each movie or music entry. The server selects content with a smaller distance from the current retrieval vector as candidate results based on vector similarity calculations.

[0430] To further enhance retrieval capabilities, the server utilizes a generative artificial intelligence model to generate internal retrieval condition suggestions, such as: "The user's current mood is 'stress,' and he likes science fiction. Please generate an internal search suggestion to recommend three science fiction movies with a slow pace, gentle visuals, and a positive ending." The server returns the internal prompt to the search module, which parses the prompt text into constraints, filtering out entries with excessive violence or depressing plots, and prioritizing content focusing on character development and themes of hope. In this way, the server no longer relies on fixed-field filtering rules, but instead uses a generative artificial intelligence model to flexibly generate structured search conditions, thereby reducing the workload of manual rule maintenance and configuration while maintaining high search accuracy.

[0431] VI. Instructions for Personality Development and Creative Assistance When a user sends a request to create a new virtual object on the terminal, the server extracts topic information from the user's input, such as: "I want to design a new fantasy character." The server constructs high-level instructions, such as: "The user requests a new fantasy character hint statement. Please output a 'personality generation instruction statement' to guide subsequent character development. The instruction should include: character identity, internal conflict, and story development strategy." The server inputs this instruction into the generative artificial intelligence model. The model then generates a structured personality generation instruction statement at the output, for example: “Design a young guardian of the kingdom who appears brave and confident, but is afraid of inheriting his father’s failed history. Let the story revolve around how he breaks free from the shadow of his family.” The server sends the creative reference information to the terminal. The terminal displays "Character Identity," "Internal Conflicts," and "Story Development Strategy" in segments on the creation interface, and provides users with separate editing areas. Users can further refine character backgrounds and personality hints based on these instructions. Through this model, the server not only provides the final content but also the meta-information of the "generation instructions themselves," essentially building a high-level creative structure for the user. This high-level instruction generation is difficult to automate in traditional manual template systems, facilitating rapid conceptualization in the early stages of creation and reducing ineffective input attempts.

[0432] VII. Technical Effects and Improvements in Computer Technology The server links natural language parsing, generative AI model inference, sentiment estimation, and content retrieval modules within a unified architecture, creating a closed-loop data flow from input, prompts, sentiment, and content recommendation. Utilizing multi-dimensional feature vectors and attention mechanisms, the server merges personality prompts with real-time sentiment states as model input conditions, thus simultaneously completing semantic generation and style control within a single model inference iteration. This holistic design differs from the traditional approach of simply using the model as a black-box text generator, achieving higher computational efficiency and finer-grained control precision at the system level.

[0433] The server continuously records user editing behavior and sentiment curves to build multi-version histories and sentiment annotations for each personality prompt. This specific data structure allows the recommendation module to reuse historical inference results in scenarios with similar emotions and similar character settings, avoiding generation from scratch and reducing the burden of repetitive inference. Simultaneously, through pre-trained and incrementally updated model parameters, the server can gradually improve generation quality while maintaining inference speed.

[0434] In this system, the terminal not only serves as a display and input device but also undertakes front-end preprocessing responsibilities. The terminal performs feature extraction or compression of images and audio locally, reducing the transmission of raw multimedia data over the network and lowering bandwidth consumption and server-side decoding burden. Since the server only receives feature vectors rather than raw streaming media data, the network protocol stack load is reduced, communication latency is lowered, and the overall interactive experience is improved.

[0435] Through the specific structural and dataflow design described above, this system's processing goes beyond simply automating human manual steps. Instead, it constructs a new computational paradigm through the collaborative work of generative artificial intelligence models and sentiment estimation models, enabling the generation, adjustment, and content recommendation of personality prompts to be technically and efficiently coupled. This design brings about multiple technical benefits, including improved generation accuracy, reduced response time, optimized utilization of computing resources, and improved storage access efficiency. It provides a directly implementable technical solution for achieving high-quality interactive content generation and intelligent recommendation in the real world.

[0436] use Figure 14 The processing procedure is explained.

[0437] Step 1: Users enter their role requirements on the terminal.

[0438] Input: Text entered by the user via the terminal keyboard or touch interface, such as "I want a brave female swordsman to fight in a future city". Output: A structured request data segment in the terminal memory, which is then sent to the server over the network as a request message.

[0439] The specific steps are as follows: The terminal encapsulates the string entered by the user in the text input box into a request object containing the user identifier, timestamp, and text content, such as fields like "user_id", "text", and "locale". Then, it uploads this request object to the server via HTTP or WebSocket protocols. Before sending, the terminal can perform basic encoding (such as UTF-8 encoding) and add simple validation information (such as length checks) to ensure the server can correctly receive and parse the text.

[0440] Step 2: The server receives and parses user input requests.

[0441] Input: A network request message containing user text information sent by the terminal.

[0442] Output: The user input text as represented internally by the server, along with metadata such as the user identifier and session identifier associated with the request.

[0443] The specific steps are as follows: After receiving a network request in the Web service module, the server's request processing thread parses the HTTP header and request body, extracts the "text" field from JSON or other formats as the raw text, and stores information such as "user_id" and "session_id" into the session management structure. The server then calls the natural language processing module to perform word segmentation, part-of-speech tagging, and syntactic analysis on the raw text, obtaining a set of labeled sequences and dependency relations, preparing clean input data for subsequent generative artificial intelligence model processing.

[0444] Step 3: The server uses natural language processing to extract semantic features from user input and construct prompts for a generative artificial intelligence model.

[0445] Input: The original user text obtained in step 2 and its intermediate results such as word segmentation and dependency structure.

[0446] Output: A prompt statement (prompt text) for a generative artificial intelligence model, such as a task instruction for generating personality prompt statements.

[0447] The specific steps are as follows: The server extracts key nouns (such as "female swordsman"), adjectives (such as "brave"), and environmental descriptions (such as "futuristic city") related to the character from the parsing results, and fills them into an internal template to construct a task description. For example: "Generating personality prompts for a character based on the following user requirements: 'A brave female swordsman fights in a futuristic city.' The personality prompts need to describe: speaking style, behavioral characteristics, values, and typical scenarios." The server uses this task description as the input sequence for the generative artificial intelligence model, and simultaneously sets model parameters (such as maximum output length, temperature, and sampling strategy) to form a complete model call request.

[0448] Step 4: The server invokes a generative artificial intelligence model to generate initial personality prompts.

[0449] Input: The prompt statement constructed in step 3 and the generated parameters.

[0450] Output: A text containing initial personality prompts, such as a complete paragraph describing the personality and behavioral patterns of the virtual object.

[0451] The specific steps are as follows: The server encodes the prompt statement into a sequence of tags, maps it to a vector through an embedding layer, and then feeds it into a multi-layer self-attention network structure. During the decoding phase, the server generates output tag by tag in an autoregressive manner until a termination tag is generated or the length limit is reached. During model inference, the server focuses on key tags such as "brave," "female swordsman," and "future city" through attention weight calculation, thereby reinforcing these concepts in the output. After inference, the server decodes the output tag sequence into a text string, forming the initial personality prompt statement, and temporarily stores this text along with the user identifier in a database table.

[0452] Step 5: The server returns the initial personality prompt to the terminal and displays it.

[0453] Input: The personality prompt text generated in step 4, along with the corresponding user identifier and prompt text identifier.

[0454] Output: The response data packet sent to the terminal, and the editable personality prompt statement displayed on the terminal screen.

[0455] The specific steps are as follows: The server constructs a response object containing fields such as "prompt_id" and "prompt_text," and returns it to the terminal via an HTTP response. Upon receiving the response, the terminal creates a text editing component in the application interface, displays the "prompt_text" content to the user, and sets the cursor position and editing event listener for each line or paragraph, allowing the user to modify, delete, or append text. The terminal also caches the prompt statement and its identifier locally for later editing and submission.

[0456] Step 6: Users can edit personality prompts on the terminal and submit the changes.

[0457] Input: The initial personality prompt text displayed in the terminal interface.

[0458] Output: The modified new personality prompt text, along with the prompt identifier, is sent to the server as an update request.

[0459] The specific action is as follows: The user modifies some content in the terminal text box, for example, adding at the end of the original text: "She always cracks a joke to ease the atmosphere after a battle, such as making fun of her own disheveled appearance." When the user clicks the "Save" or "Next" button, the terminal assembles the complete content in the current text box, the original prompt_id, and the user identifier into an update request, and sends it to the specified interface on the server using the POST method. Before sending, the terminal can check the text length and character set to avoid server-side errors caused by excessively long text or illegal characters.

[0460] Step 7: The server receives the personality prompts edited by the user and updates the storage.

[0461] Input: An update request sent by the terminal, containing prompt_id and new text content.

[0462] Output: The updated personality hint statements recorded in the database, and the latest personality hint statements for use in generating subsequent content.

[0463] The specific steps are as follows: The server parses the request body in the update interface, reading the `prompt_id` and the new `prompt_text`. The server looks up the corresponding record in the database based on the `prompt_id`, stores the old text in the version history table or history column, writes the new text to the current version column, and records the update time and editing source. The server updates the cache of prompts related to the current session in memory, ensuring that subsequent generated content references the latest text. Through this versioned update, the server can revert to historical versions when needed and perform offline evaluation of the generated results from different versions.

[0464] Step 8: The terminal collects the user's image and sound features and uploads them to the server for emotion estimation.

[0465] Input: User authorization to access the camera and microphone, as well as real-time captured video frames and audio clips.

[0466] Output: Image and audio feature vectors preprocessed by the terminal, and feature data packets uploaded to the server.

[0467] The specific steps are as follows: The terminal reads image frames from the camera at a fixed frame rate, performs preprocessing on each frame such as scaling and grayscale conversion; the terminal uses a local lightweight model or feature extraction algorithm to crop and encode the face region into feature vectors (such as keypoint coordinates or convolutional features). For audio, the terminal reads the audio stream from the microphone, calculates the short-time Fourier transform coefficients or Mel-frequency cepstral coefficients, and stacks the features from multiple frames into a time-series vector. The terminal packages these features in binary or JSON format, along with metadata such as timestamps and session identifiers, and sends them to the server through a secure channel.

[0468] Step 9: The server estimates the user's emotional state based on the characteristics of the uploaded images and sounds.

[0469] Input: Image feature vector and audio feature vector uploaded in step 8.

[0470] Output: A label representing the user's current emotional state and the corresponding probability distribution, such as "stress", "frustration", "joy", etc.

[0471] The specific steps are as follows: The server inputs image features into subsequent layers of a convolutional neural network and audio features into a recurrent neural network or a one-dimensional convolutional network. Then, the feature vectors from both modalities are connected in a fusion layer. At the output layer, the server uses a softmax function to normalize the scores for various emotions, obtaining a set of emotion probabilities. The server determines the final emotion state by selecting the emotion label with the highest probability or by using a threshold strategy, such as classifying it as "stress." The server stores the current emotion state and its probability in a session state structure and records the time series in the log to support subsequent trend analysis and model retraining.

[0472] Step 10: The server inputs the emotional state and the current personality prompt statement into the generative artificial intelligence model based on the information processing results to generate an adjusted personality prompt statement.

[0473] Input: The latest personality cue text from step 7 and the emotion status label from step 9.

[0474] Output: The new version of the personality prompt text, dynamically adjusted based on emotional state.

[0475] The specific steps are as follows: The server constructs a new prompt, merging the original personality prompt and emotion tags into the model input. For example, "Here is a personality prompt: 'This AI assistant speaks calmly and directly, often demanding that users complete tasks on time in a commanding tone.' The user's current emotion is 'stress.' Please rewrite the prompt to be gentler and more reassuring while preserving the assistant's efficient characteristics." The server sends this instruction to the generative AI model, which internally assigns higher attention weight to emotion tags such as "stress," generating a softer tone. After generation, the server receives the adjusted personality prompt, such as a version with added care and encouragement, and updates the database record, indicating the associated emotional context of that version.

[0476] Step 11: The server generates content prompts (such as dialogue or story branches) based on the adjusted personality prompts.

[0477] Input: The adjusted personality prompt text obtained in step 10, and the desired output type (e.g., dialogue phrases, story branches, etc.).

[0478] Output: A set of specific content prompts, such as dialogue spoken by virtual objects to the user or descriptions of options in an interactive story.

[0479] The specific steps are as follows: The server constructs a task prompt, such as: "Below is a character personality prompt. 'This character is a female swordsman living in a high-tech futuristic city... She always cracks a joke to ease the tension after a battle.' Based on this prompt, generate three sentences suitable for speaking to the audience during a live broadcast, each no more than 30 characters." The server inputs the adjusted personality prompts along with the task requirements into a generative AI model. The model utilizes its internal language modeling capabilities and style control to generate multiple short lines of dialogue. The server processes the generated results, performing length checks, filtering sensitive content, and other data processing. Lines that meet the criteria are stored as content prompts in the content cache and returned to the terminal via an interface for display or automatic playback.

[0480] Step 12: The server generates internal search condition suggestions based on the user's emotional state and the words contained in the personality prompts, and then performs content search recommendations.

[0481] Input: the emotional state label from step 9, the personality cue text from step 7 or 10, and the content features already saved in the content storage area.

[0482] Output: Internal search criteria suggestions and a list of candidate results that match the criteria.

[0483] The specific steps are as follows: First, the server extracts keywords reflecting the theme and style from the personality prompts. These keywords are then combined with emotion tags to form the input, constructing a task instruction, such as: "The user's current emotion is 'stress,' and they like science fiction. Please generate an internal search prompt to recommend three science fiction movies with a slow pace, gentle visuals, and a positive ending." The generative AI model outputs a structured search description. The server then parses this description into specific search constraints (e.g., type = science fiction, pace = slow, emotion = warm, ending = positive) and retrieves content entries that meet these conditions from the content storage area. The server performs similarity ranking and deduplication on the search results to obtain a list of candidate content, which is then returned to the terminal as a recommendation for display to the user.

[0484] Step 13: Users receive content prompts and recommendations on the terminal and then interact with or create content.

[0485] Input: A set of content suggestion statements returned by the server and a list of candidate recommended content.

[0486] Output: Subsequent user actions, such as selecting a recommended content, using a certain statement, or continuing to edit personality prompts, etc.

[0487] The specific actions are as follows: The terminal displays a list of dialogue, story branches, or recommended movies and music on the screen in the form of lists or cards. Users can click on a dialogue to insert it into the script, click on a recommended movie to enter the playback page, or adjust character settings based on the displayed content. The terminal records these choices and sends the new behavioral data and text input to the server in the next round of interaction, allowing the server to continuously optimize the subsequent generated results. Through this series of round-trip processes, the server, terminal, and user form a loop, gradually refining the personality prompts and content selections to achieve a highly relevant and personalized interactive experience.

[0488] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0489] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0490] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.

[0491] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0492] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.

[0493] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.

[0494] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.

[0495] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0496] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.

[0497] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0498] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0499] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0500] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0501] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0502] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0503] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.

[0504] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0505] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0506] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0507] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0508] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0509] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0510] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0511] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.

[0512] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0513] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.

[0514] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.

[0515] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.

[0516] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0517] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.

[0518] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0519] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0520] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0521] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0522] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0523] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0524] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.

[0525] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".

[0526] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0527] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0528] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0529] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0530] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0531] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0532] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.

[0533] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12 to analyze the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12 to generate a menu using a generation AI. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12 to provide the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0534] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.

[0535] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.

[0536] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.

[0537] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0538] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.

[0539] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0540] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by a perspective equivalent to the field of vision of an average healthy person).

[0541] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0542] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0543] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0544] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0545] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0546] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.

[0547] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".

[0548] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0549] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0550] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0551] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0552] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0553] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0554] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.

[0555] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0556] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.

[0557] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see [reference]). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.

[0558] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.

[0559] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.

[0560] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).

[0561] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.

[0562] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."

[0563] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values ​​representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.

[0564] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).

[0565] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.

[0566] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0567] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.

[0568] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.

[0569] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.

[0570] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.

[0571] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.

[0572] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.

[0573] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.

[0574] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.

[0575] In addition, the following notes are provided in response to the above explanation.

[0576] Example 1 (Note 1) An information processing system, characterized in that it comprises: A preprocessing device for receiving input information, including character attribute information and behavior information, from a terminal, and for standardizing the input information and classifying it according to attribute items and behavior items to generate structured data; A generation apparatus for generating natural language instruction text based on the structured data, the natural language instruction text including system-level prompts and user-level prompts for a generative artificial intelligence model, and combining the system-level prompts and user-level prompts to form text for model input; A prompt generation device is used to input the model input into a generative artificial intelligence model using text, and to automatically generate prompt statements including personality prompt statements that specify the value standards of the person, behavioral rules, response methods and language styles based on the output of the generative artificial intelligence model. A post-processing device for performing content detection, deleting invalid information and formatting on the generated prompt statement, and a communication device for sending the post-processed prompt statement to the terminal. An update device for receiving user editing information or regeneration instructions for the prompt statement from a terminal, and for regenerating or updating the prompt statement based on the input information and the editing information; A learning device for learning the conditions for generating prompt statements based on user operation history information and selection history information obtained from the terminal, and reflecting the conditions for generating prompt statements in subsequent prompt statement generation to output personalized prompt statements optimized for each user. An adjustment device for dynamically changing the content or level of detail of the generated prompt statement based on status information about the user's emotional state obtained from the terminal.

[0577] (Note 2) The information processing system according to Appendix 1 is characterized in that, The preprocessing device is configured to extract multiple elements representing personality traits, social roles, activity environments, and behavioral tendencies from the input information, and to generate multiple candidate personality prompts using the multiple elements. The system is configured to receive a selection operation on at least one of the multiple candidate personality prompts via a terminal to provide the user with reference information when conceiving new personality prompts.

[0578] (Note 3) The information processing system according to Appendix 1 is characterized in that, The system includes a storage device for storing the generated prompt statements and their corresponding character attribute information, behavior information, and user operation history information. Based on the stored information, it generates statistical information categorized by character attributes and recommended templates for generating prompt statements. When generating the input text for the model, the statistical information and the recommended templates are used to improve the generation efficiency and consistency of prompt statements indicating character behavior and reactions.

[0579] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for generating prompt statements for input into a generative artificial intelligence model based on object attribute information from a terminal; A means for sending the generated prompt statement to the terminal and updating the prompt statement according to editing or additions made by the user via the terminal; A device for dynamically generating story information using the generative artificial intelligence model based on the updated prompt statements and the object attribute information; A device for sending the generated story information to the terminal in real time, so that the terminal can display the story information sequentially; An apparatus for repeatedly executing the prompt statements and generating the story information based on new attribute information from the terminal, so as to provide different story developments.

[0580] (Note 2) The information processing system according to Appendix 1 is characterized in that, The prompting statement is configured as a personality prompt containing instructions about the object's behavior and reactions.

[0581] (Note 3) The information processing system according to Appendix 1 is characterized in that, The information processing device is configured to use the generative artificial intelligence model to generate information for use as a reference for the compilation or editing of the personality prompts, and to provide the information to the terminal.

[0582] Example 2 (Note 1) An information processing system, characterized in that it comprises: An apparatus for using a generative artificial intelligence model to perform natural language processing parsing on information obtained from a user, and for automatically generating an initial scheme of prompt statements as input to the generative artificial intelligence model based on the parsing results; A device for presenting the initial scheme of the prompt statement on the user interface of a communication terminal with display and input functions, and for receiving text information input by the user to add, delete or modify the prompt statement, thereby enabling the prompt statement to be customized by the user. A device for storing user-customized prompts in association with version information in a storage device, and for evaluating whether the customized prompts are effective input to a generative artificial intelligence model, inputting the customized prompts into the generative artificial intelligence model or a generative artificial intelligence model for evaluation that is different from the generative artificial intelligence model, and having the generative artificial intelligence model for evaluation generate feedback information including the advantages, improvements and revision schemes of the prompts; A device for presenting the feedback information on the user interface and enabling the user to repeatedly edit the prompt statement based on the feedback information, thereby refining the prompt statement based on multi-stage interaction; An apparatus for receiving a final prompt selected by a user from a plurality of stored versions of prompts, and outputting the final prompt as input data to an external generative artificial intelligence model or an external information processing device.

[0583] (Note 2) The information processing system according to Appendix 1 is characterized in that, The device for presenting feedback information is configured to display the revision schemes contained in the feedback information as candidate prompt statements on the user interface, and when responding to the user's selection operation of the candidate prompt statements, automatically reflect the selected candidate prompt statements as customized prompt statements in the prompt statements of the editing object to assist the user in correcting the prompt statements.

[0584] (Note 3) The information processing system according to Appendix 1 is characterized in that, The system is configured to store the editing history of the prompt statements as historical information in a storage area within the communication terminal or in a storage device connected via a network, and to retrieve past versions of the prompt statements in response to a user's history viewing operation, presenting the past versions as input candidates or editing candidates to be re-inputted into the generative artificial intelligence model, thereby enabling the user to compare and select among multiple versions to design prompt statements.

[0585] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: The information processing unit in the information processing device is configured to: use a generative artificial intelligence model to perform natural language processing parsing on user input information obtained from a communication terminal, and automatically generate personality prompt statements to instruct the behavior or reaction of virtual objects based on the user input information; The automatically generated personality prompt statement is output to the display area of ​​the communication terminal, the user's editing or appending operation content on the string in the display area is obtained, and the personality prompt statement is updated based on the obtained operation content, so that the personality prompt statement can be personalized by the user. Emotion estimation processing is performed based on image or sound information sent by the communication terminal to obtain the user's emotional state in chronological order. The acquired emotional state and the personality prompt statement are used as input to execute the generative artificial intelligence model again, generating an adjusted personality prompt statement that dynamically changes the content or tone of the personality prompt statement according to the emotional state. Based on the adjusted personality prompts, content prompts are generated, including virtual object dialogues, behavioral descriptions, or interactive story branch information, and the content prompts are sent to the communication terminal to provide users with interactive information presentation or creative assistance. Furthermore, the recommendation information generation unit uses the user's emotional state and the words contained in the personality prompt statement as features to search the content storage area, extract information content that matches the emotional state from the content storage area, and present it to the communication terminal.

[0586] (Note 2) The information processing system according to Appendix 1 is characterized in that, The information processing unit is further configured to: enable the generative artificial intelligence model to generate internal retrieval condition prompts related to the user's emotional state and the personality prompts, and extract candidate content from the content storage area based on the retrieval condition prompts, thereby having the recommendation information generation unit perform the extraction of information content.

[0587] (Note 3) The information processing system according to Appendix 1 is characterized in that, The information processing unit is further configured to: when receiving a new virtual object creation request from the user, use the generative artificial intelligence model to generate personality generation instruction statements containing the virtual object's role, internal contradictions, and story development guidelines as creative reference information, and present the creative reference information to the communication terminal to assist the user in creating new personality prompt statements.

Claims

1. An information processing system, characterized in that, include: processor, The processor is configured to: Generative artificial intelligence models are used to analyze information from users and automatically generate personality cues to indicate specific behaviors or responses. The generated personality prompts can be edited or appended by the user through the interface, thus allowing for customization of the personality prompts; The system monitors the user's emotional state in real time and dynamically adjusts the personality cues based on the emotional state.

2. The information processing system according to claim 1, characterized in that, The processor is configured to use a generative artificial intelligence model to generate cues that indicate the behavior or reaction of a specific character.

3. The information processing system according to claim 1, characterized in that, The processor is configured to provide users with reference information when they conceive new personality cues using a generative artificial intelligence model.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A