Information processing system

CN122797488APending Publication Date: 2026-09-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610327042.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-19
Filing Date
2026-03-17
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0003]本发明要解决的课题在于:现有基于生成式人工智能模型的故事或绘本生成系统,通常仅根据用户输入的体裁(如冒险、幻想等)和年龄段,利用预设模板或简单规则生成相对静态的故事内容,缺乏对用户在阅读过程中的实时情绪反应的感知与反馈能力

Benefits of technology

服务器可以将作品结构数据、生成条件、生成历史存储在关系型数据库中。服务器可以设计规范化的表结构,例如:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122797488A_ABST
    Figure CN122797488A_ABST
Patent Text Reader

Abstract

The application provides an information processing system. An information processing system comprises a processor configured to: provide a user interface for specifying a genre and an age to a user; automatically generate a prompt text for instructing a generative artificial intelligence model to perform story generation based on the genre and the age specified by the user; and parse the generated story by using an emotion engine for recognizing emotions of the user, and adjust the story according to the emotions obtained by the parsing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology disclosed herein relates to an information processing system. Background Technology

[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot response to the user's speech.

[0003] The problem this invention aims to solve is that existing story or picture book generation systems based on generative artificial intelligence models typically generate relatively static story content using preset templates or simple rules based solely on the user's input genre (such as adventure, fantasy, etc.) and age group, lacking the ability to perceive and respond to the user's real-time emotional reactions during the reading process. Specifically, existing technologies have the following problems: First, the system cannot automatically identify the user's emotional state while reading different story plots, resulting in a mismatch between the generated content and the user's current psychological needs, reducing immersion and participation; Second, the generated stories are basically fixed in structure and content, and even with multiple readings, the experience gained by the same user under different emotional states is relatively similar, making it difficult to achieve personalized and dynamic story evolution; Third, when selecting story elements such as story themes, character settings, and plot development, the system relies solely on static parameters such as genre and age, without taking the user's emotional characteristics into comprehensive consideration, making it impossible to adjust the story atmosphere and emotional direction in a targeted manner, and failing to meet the needs of children and other users for emotional resonance and psychological comfort. Therefore, there is an urgent need for a system that can combine user genre and age preferences and automatically adjust story content based on user emotions in order to achieve a higher degree of personalization, interactivity and emotionally adaptive story generation. Summary of the Invention

[0004] To address the aforementioned issues, this invention proposes an information processing system, including a processor configured to: provide a user interface for specifying genre and age, enabling the user to easily input their preferred story type and applicable age range; automatically generate prompt text based on the user-specified genre and age to instruct a generative artificial intelligence model to generate story content, thereby causing the generative artificial intelligence model to generate initial story content matching the specified genre and age; and analyze the generated story using an emotion engine for recognizing user emotions, identifying the user's current emotional state by analyzing behavioral data, tone of voice, facial expressions, or physiological signals during the reading process, and adjusting the story based on the analyzed emotions. Through this configuration, the system of this invention can further enable the emotion engine to reconstruct the story elements of the generated story, allowing the generated story to dynamically change according to the user's emotions, for example, by adjusting the tension of the plot, changing the emotional tone of the ending, adding or removing companion characters or comforting plot elements, thus achieving real-time adaptive content. Furthermore, the processor is configured to automatically select the theme and characters of a picture book based on the user-specified genre and age, as well as the identified emotions. This ensures that the theme and character settings more suitable for the user's current emotional state are reflected in the story generated by the generative artificial intelligence model, thereby achieving a targeted response to the user's emotions at both the theme and character design levels. Through these technical means, the present invention achieves comprehensive utilization of genre, age parameters, and user emotional states, enabling dynamic adjustments to the structure, plot, and emotional expression of the generated story content. This effectively improves the personalization and emotional adaptability of the story generation system, solving the technical problem of the lack of emotional interaction and dynamic change in story content in existing technologies.

[0005] "System" refers to an overall device including at least one processor and optional memory, communication interface and user interface, which is a hardware and software integrated entity used to perform functions such as story generation, emotion analysis and content adjustment.

[0006] A “processor” refers to a central processing unit or equivalent device used to execute computer instructions, run generative artificial intelligence model calling programs, emotion analysis programs, and story generation and adjustment programs. It can be a single chip or a combination of multiple processing units.

[0007] "User interface" refers to the human-computer interaction interface provided by the system to the user, which receives user input information such as genre and age and outputs generated story content to the user. It includes any one or a combination of graphical user interface, touch interface, and voice interface.

[0008] "Genre" refers to the story content category that users can choose, which is used to limit the overall style and theme of the story, such as adventure, fantasy, fairy tale, science fiction, etc.

[0009] "Age" refers to a parameter used to represent the age range of the target audience. The system adapts the story's language difficulty, length, plot complexity, and picture book presentation style based on this parameter to meet the comprehension and interest needs of users of different ages.

[0010] "Prompt text" refers to natural language descriptions or structured texts that are automatically generated by the processor based on the genre and age specified by the user and used to instruct generative AI models to perform story generation. It is used to convey the story style, content requirements, and constraints to the model.

[0011] "Generative AI models" refer to AI models that can automatically generate story content based on input prompt text, including but not limited to large language models and multimodal generative models, used to output stories in text form or text-image combination form.

[0012] "Story" refers to narrative content with plot structure and character settings generated by generative artificial intelligence models based on prompt text. It can be a pure text story, a picture book script, or composite content that includes image descriptions.

[0013] An "emotion engine" refers to a functional module or program component used to identify and analyze a user's emotional state. It analyzes the user's behavioral data, voice, images, physiological signals, or interactive feedback to output emotional characteristics or emotional tags that represent the user's current emotions, and uses these to adjust and reconstruct the story content.

[0014] "User emotions" refer to the emotional state that users exhibit while reading or experiencing a story, such as happiness, fear, tension, relaxation, boredom, etc., which can be identified by an emotion engine and expressed in qualitative or quantitative form.

[0015] "Analyzing the generated story" refers to the process by which the emotion engine and / or processor analyzes the story content output by the generative artificial intelligence model, including identifying the plot structure, character relationships, emotional trends, and event developments in the story, so that adjustments or reconstructions can be made based on the user's emotions.

[0016] "Adjusting the story based on emotions" refers to the process by which the system modifies the generated story content after recognizing the user's emotional state. This includes, but is not limited to, changing the pace of plot development, the intensity of event conflicts, the narrative tone, the type of ending, and the content of dialogue, so that the story better matches the user's current emotional needs.

[0017] "Story elements" refer to the basic components that make up a story, including story theme, characters, setting, plot segments, dialogue content, narrative perspective, etc., and are used to describe the structure and content of the story.

[0018] "Reconstructing story elements" refers to the process of reselecting, replacing, adding, deleting, or rearranging existing story elements based on user emotions while maintaining the overall coherence of the story. This results in changes to the structure or details of the generated story, thereby achieving dynamic adaptation.

[0019] "Dynamic change" means that the generated story content can be automatically adjusted over time or according to changes in the user's emotional state. Users can obtain different story presentations or details in different emotional states or through multiple readings.

[0020] The "theme of a picture book" refers to the core idea or emotional thread that governs the entire story, such as "growing up bravely," "friendship and mutual assistance," and "overcoming fear," and is used to guide the overall design of the plot and the style of the illustrations.

[0021] "Character" refers to a person, animal, or anthropomorphic object that appears in a story or picture book, including protagonists, supporting characters, and other entities that participate in the development of the plot. Their personality traits, appearance, and behavior can be set or adjusted according to the genre, age, and user emotions.

[0022] "Reflecting in the story" means injecting or embedding automatically selected picture book themes and character information into the story text or story structure generated by a generative artificial intelligence model, so that the story reflects these themes and character settings in its content, description, and plot arrangement. Attached Figure Description

[0023] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.

[0024] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.

[0025] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.

[0026] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.

[0027] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.

[0028] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing device and head-mounted terminal according to the third embodiment.

[0029] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.

[0030] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.

[0031] Figure 9 This represents an emotion map that maps multiple emotions.

[0032] Figure 10 This represents an emotion map that maps multiple emotions.

[0033] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.

[0034] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.

[0035] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.

[0036] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation

[0037] Hereinafter, an example of an implementation of the system according to the present disclosure will be described with reference to the accompanying drawings.

[0038] First, let me explain the terminology used in the following instructions.

[0039] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.

[0040] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.

[0041] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.

[0042] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.

[0043] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.

[0044] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.

[0045] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.

[0046] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0047] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.

[0048] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.

[0049] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0050] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.

[0051] Figure 2The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.

[0052] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0053] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0054] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.

[0055] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.

[0056] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."

[0057] Against the backdrop of the rapid development of content generation technology based on generative artificial intelligence models, traditional story generation systems typically generate fixed text only once based on the user's input topic or keywords. These systems suffer from the following technical shortcomings in their computer implementation: (1) The generation process on the server side lacks structured utilization of user attribute information, especially the joint modeling of content category information and target age group information is relatively crude. As a result, when the generation request is passed to the generative artificial intelligence model, it only appears in the form of simple text instructions. It is impossible to finely control the input conditions of the model at the system level, thus limiting the controllability and repeatability of the server's reasoning behavior of the generative artificial intelligence model.

[0058] (2) Existing systems typically lack feedback loops that address user emotional states. After receiving the generated results, the server simply forwards them, lacking the ability to perform machine-readable correspondence and computational processing between the generated text and the user's emotional state. The computer system cannot automatically reconstruct the structural elements of the story at the internal representation layer, resulting in the system being unable to perform secondary processing and optimization of the generated content based on dynamic emotional information.

[0059] (3) The technical solution for the content generation system to control the input of generative artificial intelligence models under multi-dimensional constraints (such as content category, age group, emotional state, narrative style, etc.) is not perfect. The server lacks a mechanism to map multiple high-level semantic attributes into a unified prompt statement structure. It is impossible to organize, combine and attach these attribute information in the computer with a unified data structure and processing flow, thereby reducing the scalability and reusability of the system in different task scenarios.

[0060] (4) In practical application scenarios, users of different age groups have different requirements for text length, expression complexity and plot stimulation. Traditional systems often rely on simple truncation or rule filtering after generation to "prune" the text, without systematically encoding age groups and emotional constraints in the construction stage of prompt statements, resulting in insufficient server efficiency in computing resource utilization and generation quality control.

[0061] Therefore, it is necessary to provide a new computer implementation scheme that obtains users' content category information and target age group information in a structured manner on the server side, combines the sentiment analysis results of the story text, automatically generates prompts and attribute information with a unified format, and interacts with the generative artificial intelligence model based on the prompts. At the same time, after the model outputs, the story components are re-edited and reconstructed in a programmatic way, thereby improving the calling process, input control method and dynamic adjustment capability of the output content of the generative artificial intelligence model at the computer technology level.

[0062] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.

[0063] In this invention, the server includes: an acquisition unit configured to generate an interface screen for user operation via a display device, enabling the user to input content category information and target age group information, and to acquire the selection information; an acquisition unit configured to receive communication data containing the selection information from an information terminal, and dynamically generate a prompt statement according to the type and value of the selection information and a predefined statement structure to instruct a generative artificial intelligence model on the conditions for generating content, and send generation request data containing the prompt statement to the generative artificial intelligence model; an adjustment unit configured to acquire content data generated in response to the prompt statement from the generative artificial intelligence model, and parse the expressive content contained in the content data using an emotion recognition processing unit, correlate the parsing result with the user's emotional state, and modify or add at least a portion of the constituent elements contained in the content data based on the emotional state to generate adjusted content data; and an output unit configured to convert the adjusted content data into output data in an output format that can be displayed on the information terminal, and send the output data to the information terminal via a communication link. This allows for the formation of a closed-loop data processing chain within the computer, unifying the user's content category information, target age group information, and emotional state into standardized prompts and attribute information. This enables fine-grained control over the input to the generative AI model, and dynamic adjustment of the story content through programmatic reconstruction and re-editing after the model outputs. Consequently, it improves the server's controllability, scalability, and adaptability to different user groups and emotional states during the generation process, achieving technical improvements in the way generative AI models are invoked and the quality of content generation.

[0064] "System" refers to the whole consisting of information processing devices, information terminals and their running program modules, communication links and data storage resources, used to realize a computer implementation scheme for story generation and output control based on generative artificial intelligence models.

[0065] "Information processing device" refers to an electronic computing device with a processor and memory, which can be a server or other general-purpose computing device, used to execute program instructions and process user information, prompts, and data.

[0066] "Information terminal" refers to a user-side computing device connected to an information processing device via a communication network, including but not limited to mobile terminals, tablet terminals, or display terminals, used to present interface screens and display content.

[0067] "Display device" refers to an output device connected to an information processing device or information terminal, used to visually present interface images and content to the user, including but not limited to a display screen, monitor, or touch screen.

[0068] A "processing unit" refers to a logical collection of software modules or functional modules executed by a processor on an information processing device, used to realize various processing functions such as acquiring, generating, parsing, adjusting and outputting information data.

[0069] The “acquisition unit” refers to a functional module implemented on an information processing device, which is configured to generate an interface screen for user operation, receive selection information from the information terminal, and record content category information and target age group information in the internal storage structure.

[0070] The “sending unit” refers to a functional module implemented on the information processing device, which is configured to dynamically generate prompt statements based on the selection information and send the generation request data containing the prompt statements to the generative artificial intelligence model through the communication interface.

[0071] The "adjustment unit" refers to a functional module implemented on the information processing device, which is configured to acquire narrative data generated by a generative artificial intelligence model, and modify or add elements to the narrative data based on the emotional state obtained by the emotion recognition processing unit, thereby generating adjusted narrative data.

[0072] "Output unit" refers to a functional module implemented on an information processing device, configured to convert adjusted text data into an output format suitable for display on an information terminal, and send the output data to the information terminal via a communication link.

[0073] "Content Category Information" refers to the classification information used to indicate the content type of the story to be generated, including but not limited to themes such as adventure, fantasy, and bedtime stories.

[0074] "Target age group information" refers to attribute information used to indicate the age range of users to whom the story content is suitable, including but not limited to a specific age or age group, used to constrain language difficulty and plot complexity.

[0075] "Selection information" refers to the combination of content category information and target age group information entered by the user through the interface screen and sent to the information processing device via the information terminal.

[0076] "Communication data" refers to data units transmitted through communication networks between information processing devices and information terminals or between information processing devices and generative artificial intelligence models, including request data and response data.

[0077] "Generative artificial intelligence models" refer to models trained based on machine learning algorithms, especially deep learning and natural language processing techniques, that can automatically generate content such as text based on input prompts.

[0078] "Prompt statements" refer to text instructions generated by an information processing device according to a predefined statement structure based on selected and attribute information, used to explicitly instruct generative artificial intelligence models on the conditions for generating narrative content.

[0079] "Generate request data" refers to a request message structure that includes prompt statements and parameters related to model generation, used to request the generative artificial intelligence model to perform story generation processing.

[0080] “Story data” refers to the text data output by generative artificial intelligence models based on prompts, representing story content with a plot structure, which may include a beginning, development, and ending.

[0081] "Emotion Recognition Processing Unit" refers to a functional module or external service implemented on an information processing device, used to perform emotion analysis on text data or emotion recognition on user response data in order to obtain information about the user's emotional state.

[0082] "Emotional state" refers to the state information identified by the emotion recognition and processing department that represents the user's emotions or emotional tendencies, including but not limited to categories such as happy, nervous, and calm, which is used to guide the adjustment of the story data.

[0083] "Constituent elements" refer to the components that make up the story data, including but not limited to the main characters, scene settings, event progression, dialogue content, and ending structure.

[0084] "Reconstruction conditions" refer to a set of rules or parameters determined by an information processing device based on information such as emotional state, used to guide the recombination, replacement, or adjustment of the constituent elements in the narrative data.

[0085] "Adjusted story data" refers to updated story text data obtained by changing or adding constituent elements based on the initial story data, according to emotional state and reconstruction conditions.

[0086] "Attribute information" refers to high-level semantic information used to limit the generation behavior of generative artificial intelligence models, including the category of the topic, the category of the subject, and the category of the expression style, etc., which are used to refine the generation conditions in the prompt statements.

[0087] "Story Theme Category" refers to classification information used to indicate the overall theme of the story, including but not limited to friendship, courage, growth, and family.

[0088] "Appearance Category" refers to the classification information used to indicate the type of main characters in the story, including but not limited to child characters, animal characters, fantasy characters, etc.

[0089] "Expression style category" refers to classification information used to indicate the language style and narrative method of a story, including but not limited to gentle narration, humorous narration, tense narration, etc.

[0090] "Storage device" means a computer-readable storage medium used to store program instructions, configuration data, attribute information and object data, including but not limited to semiconductor memory, magnetic storage medium or optical storage medium.

[0091] "Output format" refers to the data form after the adjusted story data has been organized and encoded in order to be displayed correctly on the information terminal, including but not limited to structured text format and data structure corresponding to the interface layout.

[0092] A "communication link" refers to a physical or logical connection used to transmit communication data between information processing devices, information terminals, and generative artificial intelligence models, including but not limited to wired network connections and wireless network connections.

[0093] In embodiments of the present invention, the server, terminal, and user collaboratively use a generative artificial intelligence model to perform story generation and emotion adaptive adjustment processing. The following description uses the server, terminal, and user as subjects to illustrate the specific implementation of this system.

[0094] The server can use general-purpose server hardware. It includes a central processing unit (CPU), main memory (RAM), non-volatile memory (SSD or hard disk drive), and a network interface card. For example, the server can be configured to run web server software (such as Nginx or Apache) and application servers (such as Python-based Django / Flask or Node.js-based Express) on a Linux-based operating system (such as an Ubuntu-based server). The server can deploy generative AI models on GPU servers or access generative AI model APIs provided as cloud services.

[0095] The terminal can use portable information terminals (smartphones), tablets, personal computers, and other electronic devices as hardware. The terminal includes a display device (LCD, OLED, etc.), an input device (touchscreen, keyboard, mouse, etc.), and a network communication module (Wi-Fi module, mobile communication module, wired network interface, etc.). The terminal can display HTML / JavaScript interfaces in web browsers (such as Chrome and Safari) or display a native UI as a mobile application (such as an Android-based application or an iOS-based application).

[0096] Users specify content category information (e.g., adventure, fantasy, bedtime stories) and target age group information (e.g., 3 years old, 4 years old, 5 years old) through the interface displayed on the terminal. Users select a suitable combination by touching or clicking on drop-down lists, radio buttons, and other controls displayed on the terminal display device.

[0097] To receive selection information sent from the terminal, the server uses middleware (such as Gunicorn, uWSGI) and web application frameworks (such as Django, Flask, Express) to parse network requests. The server receives communication data sent from the terminal in structured form such as JSON, and the parsing module extracts the fields "genre" (content category information) and "age" (target age group information) and stores them in internal data structures (such as key-value maps, objects).

[0098] To generate prompts based on the obtained content category and target age group information, the server uses a template storage module and a template rendering module. The server maintains multiple prompt templates on the storage device. A prompt template is natural language text containing placeholders for inserting parameters such as content category, age group, and sentiment attributes, for example, in the following form: Example prompt template 1: "Please create a story of {content category} for children of {age group}. The language should be simple, interesting, and contain positive educational value." Example 2 of prompt statement template: "Please generate a short story of {content category} for children of {age group}. The sentences should be very short, the vocabulary should be very basic, and there should be no complex plot." The server selects an appropriate template based on the chosen information and generates prompts for the generative AI model by substituting specific parameters into placeholders. For example, when a user selects "Content Category: Adventure" and "Target Age: 5 years old," the server generates the following prompt: "Please create an adventure story for a 5-year-old child, using simple, fun language and containing positive educational value." When a user selects "Content Category: Fantasy" and "Target Age Group: 3 years old", the server generates the following prompt: "Please generate a short fantasy story for a 3-year-old. The story can include talking animals and a magic tree. Sentences should be very short, vocabulary should be very basic, and complex plots should be avoided." By constructing prompts with this unified structure, the server standardizes the natural language instructions input to the generative artificial intelligence model and systematically realizes the conversion between internal data structures (content categories, age groups, and emotional attributes) and natural language prompts. The server employs a large-scale language model based on deep learning as the generative AI model. Generative AI models, for example, use a multi-layer Transformer architecture. The generative AI model lexically converts the input text (prompt statements), transforming each word into an embedding vector, and then stacks multiple layers of encoder / decoder structures containing multi-head self-attention layers and feed-forward neural network layers. During training, the generative AI model uses a large-scale text corpus and employs a cross-entropy loss function to minimize the error in the next word prediction task. Weight parameters are updated using backpropagation and optimization algorithms (such as the Adam optimizer).

[0099] During generative AI model inference, the server can explicitly specify hyperparameters (such as maximum output length `max_tokens`, temperature parameter, top-k or top-p sampling parameters). The server sends the prompt statement along with these parameters as generation request data to the model service, thereby controlling the model's generation behavior (diversity, stability, length).

[0100] To communicate with the model service, the server can use an HTTP-based REST API or a remote procedure call (such as gRPC). The server serializes the request, which includes a prompt statement, into JSON format or similar and sends it via the network interface to the model service running on the GPU server. The server receives a data structure containing the generated story text from the response returned by the model service.

[0101] The server performs sentiment recognition processing on the received story data. The server can use a classification model or a sequence labeling model as the sentiment recognition processing unit. For example, the sentiment recognition processing unit can be composed of a sentiment classification network using a Bidirectional Long Short-Term Memory (BiLSTM) network, a Convolutional Neural Network (CNN), or a lightweight Transformer model. The sentiment recognition processing unit lexically converts the story text into a vector representation through a word embedding layer, and calculates the sentiment score of the entire text (e.g., positive, neutral, negative, tense, calm, excited, etc.) through multiple neural layers. The server uses a softmax function to calculate the probability of each sentiment category on the output vector of the sentiment recognition model, and adopts the category with the highest probability as the user's sentiment state.

[0102] The server reconstructs the components of the story data based on the user's emotional state. To internally break down the story data into units such as paragraphs, sentences, main characters, scene descriptions, event progression, and ending elements, the server can use a rule-based parser or a dependency parser. The server saves the parsed results as structured data (such as a story graph composed of nodes and edges, or tree-like data representing a hierarchical structure).

[0103] The server defines the correspondence between emotional states and various elements of the story structure as a set of rules. For example, when the emotional state is judged to be "too tense," the server applies the rule to change the ending to a "warm resolution," and to reduce the number of events that cause tension, it deletes specific event nodes or replaces them with gentler descriptions. When the emotional state is judged to be "too calm to maintain attention," the server applies the rule to add adventure or surprise elements.

[0104] In this rule-based reconstruction process, the server can use a pre-designed, non-idiomatic structure processing order instead of simply replacing parts of the statements. For example, the server can process the data in the order of (1) rearranging the character list, (2) rearranging the scene order, (3) inserting conflict events, and (4) changing the ending type, inserting natural connecting sentences at each step using simple language generation templates. Thus, the server can reuse the large-scale output of generative artificial intelligence models while re-editing the story data at the component level, thereby achieving the reuse of computing resources and improving the stability of the results.

[0105] Based on content category information, target age group information, and emotional state, the server selects attribute information such as story theme category, main character category, and expression style category from the storage and appends this attribute information to the prompt statement. For example, when targeting a younger age group, the server selects "short sentences, simple vocabulary" as the expression style category; when the emotional state is unstable, the server appends "reassuring, friendly" as the story theme category. The server converts this attribute information into natural language and adds it as part of the prompt statement. For example, the server generates the following prompt statement: "Please create an adventure story for a 5-year-old child. The story should reflect the themes of friendship and cooperation. The language should be simple and fun, the sentences should not be too long, and the ending should be heartwarming and reassuring." By using these integrated multi-attribute prompts, the server can refine the input conditions for the generative AI model and constrain the attention distribution and generation direction within the model. This improves the consistency and adaptability of the generated results, and by reducing the amount of correction required for subsequent reconstruction processing, it achieves the technical effects of improved overall processing speed and reduced computational cost.

[0106] The terminal receives output data from the server and uses a JSON parsing module to extract the story text and metadata (content category, age group, etc.). The terminal uses a display layout engine to break the text into lines and adjusts font size, line spacing, background color, etc., according to age group. For example, for younger users, the terminal uses a larger font and high-contrast color scheme to improve readability. The terminal paginates the text blocks according to the display device's resolution and orientation (portrait / landscape) and displays the story using scrolling or page-turning animations.

[0107] Users read the adjusted story text on the terminal display. Users can then use the Regenerate or Settings buttons to reselect content categories or age groups, and re-attempt emotional adjustment. These user actions generate new selection information, which is then sent back to the server.

[0108] In this system architecture, the server goes beyond simply automating manual operations. It improves computer technology itself through standardized internal information representation, modularized prompt generation, and combined algorithms for sentiment analysis and story reconstruction. By templated prompts and added attribute information, the server structures the input space of the generative AI model and ensures consistency in model calls, thereby suppressing fluctuations in output across requests and reducing errors (deviations between user requests and generated content). By combining a sentiment recognition processing unit with a structured reconstruction algorithm, the server can generate stories that more closely resemble the target emotional profile than by directly using the model output, thus achieving a substantial "accuracy improvement."

[0109] The server reduces the amount of filtering and correction processing in the backend by encoding constraints into prompts based on age and content category at the model invocation stage, thereby reducing the overall number of model invocations and regenerations required. Therefore, the server-side computational and communication load is reduced, resulting in shorter response times and improved processing throughput.

[0110] In generative AI model learning methods, servers can combine pre-training and fine-tuning. During pre-training, the server trains the language model using a large-scale general text corpus, followed by fine-tuning using children's story texts and sentiment-annotated datasets. During fine-tuning, the server can add a sentiment consistency term to the loss function to minimize the difference between the target sentiment label and the sentiment estimate of the generated text, thereby updating the weights. Through this training process, the server guides the model to more sensitively reflect the age group, content category, and sentiment constraints in the prompts.

[0111] As a method to expand training data, servers can use sentence order perturbations, synonym substitutions, and sentence splitting and combination within the story text. By using these expansion techniques, servers increase the diversity of training data, thereby improving the model's generalization performance and robustness. Through such internal algorithms and data flow design, servers enhance the output quality of generative AI models and the overall stability of the system.

[0112] In addition to the implementation forms described above, the server can also adopt various variations. As a generative artificial intelligence model, the server can use architectures other than Transformer (e.g., models based on RNNs or hybrid attention mechanisms). As a sentiment recognition processing unit, the server can also use dictionary-based sentiment analysis algorithms or rule-based classifiers to refine the results of the neural network. As a method for generating prompt statements, the server can, in addition to a single template approach, employ a decision tree-based approach combining template fragments, or a method using Markov decision processes to automatically explore the structure of the prompt statements.

[0113] In terms of display methods, in addition to a single-page view, the terminal can also be combined with page-turning animations, voice reading functions, and image insertion display functions. The terminal can also be configured to use a speech synthesis engine to convert story text into speech data and output it from a speaker. In this case, the terminal combines the control of physical sound output devices by passing structured text data received from the server to the speech synthesis module, thus going beyond abstract information processing and providing real-world technology usage forms.

[0114] By using this system, users can not only input text and generate stories, but also indirectly control the generation process through parameters such as age group, content category, and emotional state. However, users do not need to be aware of the complexity of the internal algorithms. Because the server automatically performs processes such as prompt statement generation, model invocation, and emotional reconstruction based on the aforementioned data structures, algorithms, and models, users can obtain high-quality and individually tailored stories.

[0115] As described above, in the embodiments of the present invention, the server, terminal, and user each assume specific roles, forming a technical processing flow centered on a generative artificial intelligence model and prompt statements. The server achieves technical effects different from traditional simple automatic article generation systems (increased processing speed, improved generation accuracy, reduced computational load, and improved communication efficiency) through internal computer technology processing such as the integrated generation of prompt statements, the encoding of attribute information, and the reconstruction of stories based on emotional states.

[0116] use Figure 11 The processing flow is explained.

[0117] Step 1: Users select content categories and target age groups on the device.

[0118] Users can select an item from a list of content categories (such as "Adventure", "Fantasy", "Bedtime Stories") and an item from an age group list (such as "3 years old", "4 years old", "5 years old") by clicking drop-down boxes or radio buttons on the interface using the terminal's touch screen or mouse.

[0119] Input: Content category options and age group options pre-displayed on the interface.

[0120] Output: Information on the content category currently selected by the user and the target age group.

[0121] The user's selection action causes the terminal to update the corresponding state variables, storing the selected category and age in memory as key-value pairs for later packaging and sending to the server.

[0122] Step 2: The terminal sends request data containing selection information to the server.

[0123] After the user clicks the "Generate Story" button, the terminal reads the content category information and target age group information stored in memory, constructs a request object containing these two fields, and serializes it into a JSON string. The terminal then calls the network communication library and sends this JSON as the request body to the server's preset interface address via an HTTP POST request.

[0124] Input: Content category information and target age group information selected by the user.

[0125] Output: HTTP request data sent to the server over the network (including selection information in JSON format).

[0126] Before sending, the terminal sets the request header (such as Content-Type as application / json) and encapsulates the server address, interface path and request method together to form a complete network data packet.

[0127] Step 3: The server receives and parses the selection information from the terminal.

[0128] After receiving an HTTP request in the network stack, the web server and application framework work together to read the JSON string from the request body and call a JSON parsing library to convert it into an internal data structure. The server then extracts the content category field and the target age group field from the parsed structure and performs validity checks.

[0129] Input: HTTP request data containing content category information and target age group information.

[0130] Output: Structured selection information (such as internal objects or dictionaries) represented in server memory.

[0131] The server checks the field values ​​against a predefined list of allowed values. If a value is invalid, an error response is generated. If a value is valid, the structured selection information is passed to the subsequent prompt generation module.

[0132] Step 4: The server generates a prompt statement based on the selected and attribute information.

[0133] The server selects prompt templates corresponding to the content category and target age group from its internal template storage, and queries attribute information (such as story theme category, main character category, and expression style category) in the storage device when needed, combining these attributes with the selected information. The server then replaces placeholders such as "{age group}", "{content category}", "{theme}", and "{style}" with specific values ​​through string concatenation or template rendering to generate complete natural language prompts.

[0134] Input: Structured selection information (content category, target age group) and attribute information read from storage device.

[0135] Output: Natural language prompt text used to drive generative artificial intelligence models.

[0136] For example, the server generates the following prompt: "Please create an adventure story for a 5-year-old child. The story should reflect the themes of friendship and cooperation. The language should be simple and interesting, the sentences should not be too long, and the ending should be heartwarming and reassuring." Step 5: The server constructs a model request and invokes the generative artificial intelligence model.

[0137] The server places the generated prompt statement into the model request structure, and sets generation control parameters such as maximum output length, temperature parameter, and sampling strategy. The server uses an HTTP or RPC client to serialize the request structure into JSON or binary message, and sends it over the network to the inference server where the generative artificial intelligence model is deployed.

[0138] Input: Natural language prompts and control parameters related to generation.

[0139] Output: Model request data sent to the generative artificial intelligence model service.

[0140] The server records the request ID and timestamp when sending the request, which is used for subsequent response matching and performance monitoring, and registers callbacks in the network library or waits for the model service's response.

[0141] Step 6: The server receives narrative data returned by the generative artificial intelligence model.

[0142] After the model service completes inference, the server receives a response message from the network interface and uses the corresponding parsing library to extract the generated text fields. The server treats this generated text as initial story data, stores it in memory, and associates it with the original prompts and selection information, storing it in a structured form.

[0143] Input: Response data (including generated text) returned by the generative artificial intelligence model.

[0144] Output: The initial story data text and its associated metadata stored in the server's memory.

[0145] During this process, the server can also check whether the length and encoding format of the generated text meet expectations. If an anomaly is detected, it will trigger a retry or error handling logic.

[0146] Step 7: The server performs emotion recognition and emotion state determination on the story data.

[0147] The server inputs the initial story data into the emotion recognition processing unit, which segments the text, embeds it into vectors, and then uses a pre-trained neural network (such as BiLSTM, CNN, or a small Transformer) to calculate the emotion feature vector. The emotion recognition processing unit applies a classification layer and Softmax to the output vector, calculates the probability distribution of each emotion category, and the server selects the emotion category with the highest probability as the user's target emotion state or the current emotion state of the story.

[0148] Input: Initial story data text.

[0149] Output: The emotional state label representing the emotional tendency of the story and its probability value.

[0150] The server compares the emotional state with thresholds or preset rules to determine whether the current emotional state of the story is suitable for the target age group and content category, providing a basis for subsequent editing.

[0151] Step 8: The server performs structured analysis of the elements constituting the narrative based on emotional states.

[0152] The server uses a text parsing module to segment sentences and paragraphs from the initial story data, and calls syntactic analysis tools or rule engines to extract elements such as main characters, scene descriptions, event sequence, conflicts, and outcomes from the text. The server organizes these elements into structured data, such as building a story graph containing nodes (events, characters, scenes) and edges (sequence relationships, causal relationships), or building a multi-level tree structure.

[0153] Input: Initial story data text and emotional state information.

[0154] Output: A structured representation of the story's constituent elements and their relationships (such as a story diagram or tree structure).

[0155] During the parsing process, the server uses sentiment as an additional signal to label each sentence or paragraph with sentiment intensity, which is then used to identify parts with excessively strong or weak emotions.

[0156] Step 9: The server reconstructs and adjusts the story based on emotional states and rules.

[0157] The server inputs the structured narrative representation and emotional state into the rule engine or reconstruction module. Based on predefined reconstruction conditions (such as "reduce conflict events if tension is too high" or "add mild adventure elements if emotions are too bland"), the server selects the nodes and edges to be modified. The server performs operations on the target nodes, including deleting certain events, replacing certain descriptive sentences, adding mitigating elements, or changing the ending type, and inserting connecting sentences where necessary to maintain text coherence.

[0158] Input: Structured story representation, emotional state, and set of reconstruction condition rules.

[0159] Output: The structured and adjusted story representation and the corresponding new story text (adjusted story data).

[0160] The server uses this rule-based and structural reconstruction to make fine-grained modifications to the story, thereby maintaining the overall story framework while making the emotional curve suitable for specific age groups and thematic requirements.

[0161] Step 10: The server generates output data for display on the terminal.

[0162] The server packages the adjusted narrative text and related metadata (such as content category, age group, and sentiment tags) into an output data structure and encodes it according to the terminal's expected format (such as JSON fields "content", "genre", "age", etc.). The server divides the text into paragraphs adapted to the terminal's screen as needed and reserves pagination or scrolling information for the terminal.

[0163] Input: Adjusted story data text and its associated metadata.

[0164] Output: Output data structure that can be directly parsed and displayed by the terminal.

[0165] When generating output data, the server can also record statistical information such as content length and generation time, so as to optimize model parameters and server resource configuration in the future.

[0166] Step 11: The server sends output data to the terminal via the network.

[0167] The server invokes the HTTP response interface, sending the output data as the response body to the terminal that previously made the request, and sets the content type and character encoding in the response header. After sending, the server updates its log records, including the request identifier, response time, and returned data size, to support performance monitoring and error tracing.

[0168] Input: Output data structures used for the response and terminal connection information.

[0169] Output: The HTTP response message transmitted to the terminal over the network.

[0170] The server employs compression or chunking transmission strategies when necessary to reduce communication load and improve data transmission efficiency.

[0171] Step 12: The terminal receives the server's response and parses the generated story content.

[0172] The terminal receives the HTTP response in the network library callback, reads the data in the response body, and calls the JSON parsing module to extract information such as story text, content category, and age group. The terminal stores the parsed story text in a display buffer in memory for use by the interface rendering module.

[0173] Input: The HTTP response data returned by the server.

[0174] Output: The story text and metadata available for display in the terminal memory.

[0175] During the parsing process, the terminal also checks the integrity of the fields and the correctness of the encoding. If any abnormality is found, it will prompt the user or try to make a new request.

[0176] Step 13: The terminal displays the story content on the display device for the user to read.

[0177] The terminal passes the text to the interface rendering module, setting appropriate font size, line spacing, and color scheme based on the target age group. The terminal then segments and typeset the text, rendering it on the display device and allowing users to scroll or turn pages by swiping or clicking.

[0178] Input: Parsed story text and display parameters (such as font style and layout information).

[0179] Output: A story interface that is visualized on a display device.

[0180] During the display process, the terminal can dynamically adjust the layout to adapt to different screen sizes and orientations, and display content category and age group labels when needed to help users understand the setting of the current story.

[0181] Step 14: Users read the story and interact with it on the device.

[0182] Users read the story text displayed on their visual perception terminal by scrolling or turning pages at their own pace. After the story ends, users can choose to regenerate the story, change the content category, or adjust the target age group based on their experience, triggering the terminal to send a new request to the server.

[0183] Input: The story interface displayed on the terminal and the interactive controls (buttons, menus, etc.) provided by the terminal.

[0184] Output: New user operation signals (such as regeneration command, category switching command, age group adjustment command).

[0185] These subsequent user actions will generate new selection information, which will be repeated through the terminal and server to achieve a continuous and adaptive story generation experience.

[0186] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0187] In existing content generation systems, the user-selected subject category or simple age parameter is typically used directly as input to the generative AI model without fine-grained structured design of input prompts or streaming and adaptation processing of the generated content. This results in the following shortcomings from a computer technology perspective: (1) On the input side, the computer system lacks a mechanism for programmatic combination and constraint encoding of multi-dimensional parameters such as "information category" and "user attributes". The generated prompts cannot fully express the structural requirements of the target content (such as word count constraints, expression style constraints, plot progression constraints, etc.), resulting in the inefficient use of computational resources of the generative artificial intelligence model and unstable output content quality.

[0188] (2) On the output side, computer systems typically receive and return generated content in a batch manner. They lack the technical means to manage and transmit the successive outputs of generative artificial intelligence models in segments. The terminal cannot start presenting partial results before the model calculation is finished, resulting in long waiting times for users. At the same time, the server and terminal bear a heavy burden in terms of network and memory management.

[0189] (3) On the adaptation side, the computer system lacks an algorithmic process for automatically parsing, extracting elements and dynamically adjusting the generated content based on "user attributes". It is impossible to automatically control the expression difficulty, expression length and semantic elements of the generated content on the server side through unified string parsing and symbol sequence parsing. It is difficult to achieve detailed adaptation for users of different age groups or different comprehension abilities through programming.

[0190] (4) In terms of overall architecture, there is a lack of a systematic computer implementation scheme that organically links the processing of "prompt statement generation", "streaming reception and distribution", "content post-processing adaptation" and "terminal sequential display control", so that the resource consumption and response performance of the three links of model calling, network transmission and interface presentation cannot be effectively coordinated and optimized in the end-to-end data path.

[0191] Therefore, it is necessary to propose a new system and its computer implementation scheme, which enables the server to perform structured encoding of user input, automatically generate prompts with constraints, utilize the streaming output capability of generative artificial intelligence models, perform segmented management and content adaptation processing on the server side, and realize the sequential parsing and display control of segmented information on the terminal side, thereby improving the efficiency, adaptability and interactive response performance of content generation at the computer technology level.

[0192] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.

[0193] In this invention, the server includes: an acquisition unit, which acquires information categories and user attributes specified by the user from a terminal via an information input / output device, and stores the information categories and user attributes in structured data form in computer memory; a prompt statement generation and sending unit, which automatically generates prompt statements to instruct a generative artificial intelligence model to perform content generation based on the information categories and user attributes, embedding at least one constraint information among word count conditions, expression style conditions, and / or plot progression conditions in the prompt statements, and sending the prompt statements as a character sequence to the generative artificial intelligence model; and a streaming unit, which receives sequential outputs of the prompt statements from the generative artificial intelligence model. The system generates content by dividing it into segments and caching them sequentially, and then streaming these segments to the terminal. A display control unit parses the streamed segmented information received on the terminal and appends it sequentially to a visual display device to gradually present the generated content to the user. A content adaptation unit performs string parsing and / or symbol sequence parsing on the generated content to extract content elements, determines the expression difficulty and / or expression length based on user attributes, and replaces and / or deletes at least some content elements in the prompt statements and / or the generated content accordingly, thereby obtaining adapted generated content. This allows for the programmatic encoding of user input on the server side to generate structured prompts. By leveraging the streaming output capabilities of generative artificial intelligence models, the generated content can be segmented and streamed over the network. Combined with user-attribute-based automatic content adaptation and terminal sequential display control, this improves resource utilization efficiency, response speed, and multi-user adaptability in the content generation and presentation process at the computer technology level, thereby achieving a technological improvement over existing content generation systems.

[0194] "Information input / output device" refers to a device used for information interaction between a user and a system, including various hardware and / or software components used to receive user operation input and output interface information to the user, such as touch screen devices, display terminals, human-computer interaction interface programs, etc.

[0195] "Terminal" refers to an electronic device that is directly operated by the user and communicates data with the server, including but not limited to mobile communication devices, portable computing devices, head-mounted display devices, etc.

[0196] "User attributes" refer to parameter information that characterizes user features and is used to determine the adaptation method of generated content, including but not limited to age information, comprehension ability information, and preference information.

[0197] "Information category" refers to the abstract classification of content to be generated or provided according to its theme, genre, or purpose, including but not limited to adventure, science fiction, and bedtime stories.

[0198] "Generative artificial intelligence models" refer to data processing models built on machine learning algorithms that can automatically generate text and other content based on input prompts, including language generation models obtained through large-scale training.

[0199] "Prompt statements" refer to natural language or symbol sequences used to instruct generative artificial intelligence models to generate target content, including character information that constrains or explains the content's theme, style, length, structure, etc.

[0200] The “acquisition unit” refers to a functional component in the server that receives and reads input data such as information categories and user attributes from the terminal by executing program instructions. It can be implemented by software modules executed by the processor and / or related hardware circuits.

[0201] The "prompt statement generation and sending unit" refers to a functional component in the server that automatically constructs prompt statements based on the acquired information category and user attributes, and sends the prompt statements to the generative artificial intelligence model. It can be implemented through string processing programs and communication programs.

[0202] A “streaming unit” refers to a functional component in a server that receives the generated content output sequentially by a generative artificial intelligence model, divides the generated content into segments, and continuously streams it to the terminal via the network.

[0203] "Segmented information" refers to multiple content fragments obtained by breaking down the overall content output by a generative artificial intelligence model in terms of time or structure, with each fragment being transmitted or processed as an independent data unit.

[0204] "Content data" refers to complete generated content composed of multiple segments of information combined in a predetermined order, which is stored and managed in the form of data structures on servers and / or terminals.

[0205] The "display control unit" refers to the functional component in the terminal that parses the received segmented information and controls the display mode, display timing and display layout of the visual display device, so that the generated content is presented to the user in sequence.

[0206] "Visual display device" refers to a display device used to present information to a user in the form of images, including but not limited to liquid crystal displays, organic light-emitting displays, head-mounted displays, etc.

[0207] The "content adaptation unit" refers to a functional component in the server that parses and judges the generated content and modifies or adjusts it according to user attributes, so that the generated content matches the characteristics of the target user.

[0208] "String parsing" refers to the process of scanning, segmenting, and recognizing sequences of characters in generated content to extract text units such as words, sentences, or specific tags.

[0209] "Symbol sequence parsing" refers to the process of analyzing and processing a sequence of tokens representing generated content in order to identify syntactic structures, semantic units, or other content elements. These tokens can be lexical units, coded symbols, or other abstract symbols.

[0210] "Content elements" refer to the basic components that constitute the content extracted from the generated content, including but not limited to character information, scene information, event information, plot structure, keywords, etc.

[0211] "Expression difficulty" refers to a quantitative or qualitative evaluation of the generated content in terms of vocabulary complexity, sentence structure complexity, and overall comprehension difficulty, used to measure whether the content is suitable for specific user attributes.

[0212] "Expression length" refers to the length characteristics of the generated content in dimensions such as the number of characters, words, sentences, or paragraphs, and is used to control the length of the content.

[0213] "Word count condition" refers to the constraint information used in the prompt statement to limit the total number of characters or word range of the generated content.

[0214] "Expression style conditions" refer to the constraint information in the prompt statement used to limit the style, tone, or narrative method of the generated content.

[0215] "Plot progression conditions" refers to the constraints in the prompt statements used to limit the generated content in terms of plot advancement, structural segmentation, or development pace.

[0216] In one embodiment of the present invention, the server is configured to run on a computing device in a data center, and the terminal is configured to be an electronic device carried or worn by the user. The user realizes the overall function of the system through communication between the terminal and the server.

[0217] The server, in terms of hardware, includes at least one central processing unit, a semiconductor storage device, and a network interface device. The server runs an operating system (e.g., a UNIX-based server operating system) and backend service programs on its software. These backend service programs implement functional modules such as an acquisition unit, a prompt generation and sending unit, a streaming unit, and a content adaptation unit. The terminal, in terms of hardware, includes a display module (e.g., an LCD screen or head-mounted display), an input module (e.g., touch sensors, buttons, touchpads), a processing unit, and a wireless communication module. The terminal runs a terminal operating system (e.g., a mobile terminal operating system) and a graphical user interface program to implement functional modules such as a display control unit.

[0218] In this embodiment, the server uses a generative artificial intelligence model as the core of content generation. This model employs a multi-layer self-attention network structure during the training phase. The server uses a multi-layer encoder-decoder or decoder stacked transformer architecture, with each layer including a multi-head self-attention sublayer, a feedforward network sublayer, and a normalization sublayer. During training, the server converts a large amount of text data into discrete symbol sequences. User attributes and information categories are encoded in the training data using special tags or control sequences, enabling the generative artificial intelligence model to generate text matching specific age groups and themes based on the conditional information contained in the input prompts during the inference phase. The server uses a loss function as the error function during training, calculating the loss value based on the cross-entropy between the real next symbol and the model's predicted distribution. The model parameters are updated using a gradient backpropagation algorithm, combined with an adaptive learning rate optimization method to improve convergence speed. During the training phase, the server can perform data expansion processing on the training data, such as random masking, sentence shuffling, and synonym rewriting, to enhance the model's consistency and robustness under different expression methods.

[0219] After receiving information categories and user attributes from the terminal during the inference phase, the server encodes this information into a prompt statement. The server then inputs the prompt statement as a character sequence into the generative AI model. Internally, the generative AI model converts this character sequence into a label vector sequence. The server performs a forward propagation operation on this vector sequence using a multi-layer transformer structure, calculating the context representation based on attention weights at each time step and selecting the next output label according to a conditional probability distribution. The server achieves label-by-label output of generated content through this sequence generation algorithm. Temperature parameters and candidate truncation algorithms can be used in this process to control output diversity and stability.

[0220] In this embodiment, the server divides the output from the generative artificial intelligence model into segmented information. Internally, the server maintains a data structure for each user session, including a list of segmented information, a session identifier, a user attribute summary, and an information category summary. The server sequentially appends each newly arrived segmented information to this data structure and streams these segments to the terminal via the network transmission module. The server uses chunked transmission or persistent connection transmission mechanisms in the streaming transmission, allowing the terminal to begin displaying parts of the content before the server has received all the generated content. This combination of segmentation and streaming reduces peak bandwidth usage at the computer network transmission layer and avoids the resource pressure of caching entire long texts at once in terms of memory usage, thereby achieving a balance between communication and storage loads.

[0221] After receiving segmented streaming information, the terminal maintains a text buffer structure in its local memory. This buffer structure records the received segmented content in a sequential list. Each time a new segment is received, the terminal performs character-level or tag-level parsing and updates the text buffer structure based on the parsing results. After updating the buffer structure, the terminal invokes the display control logic to append the latest content to the visual display device. This append-only display method allows users to begin reading the generated content before the generation process is complete, reducing subjective waiting time and improving the interactive experience. The terminal can also display a progress indicator based on the receiving status and adjust the refresh rate according to network conditions to avoid rendering overhead from frequent redraws.

[0222] The server further processes the generated content within the content adaptation unit. First, the server performs string parsing and symbol sequence parsing on the combined content data. During this parsing process, the server divides the text into sentences, paragraphs, and logical units based on delimiters, punctuation marks, or predefined grammatical clues. Based on this, the server extracts content elements, such as character names, scene descriptions, and event descriptions. After identifying these content elements, the server evaluates the complexity of each content element based on age or comprehension ability information contained in the user's attributes. For example, the server uses sentence length and vocabulary difficulty level as features to quantify the difficulty of expression; the server uses the total number of characters, sentences, and paragraphs as features to measure the length of expression.

[0223] When the server determines that the difficulty of the expression exceeds the target user's range, it employs a rule-based replacement or simplification strategy. The replacement strategy maintains a vocabulary or phrase mapping table, replacing complex words with simpler ones and breaking long sentences down into multiple shorter sentences. The simplification strategy removes non-critical information or redundant descriptions to reduce text complexity and length. Through this automated adaptation process, the generated content meets the comprehensibility requirements of users of different ages without requiring manual editing, achieving automatic content optimization within the computer.

[0224] When generating prompts, the server encodes user attributes and information categories as conditional control information. For example, the server may include age suggestions, subject matter descriptions, and structural requirements in the prompts. Typical examples include: "Please generate an adventure story for a 5-year-old," "Please generate a science fiction story for a 10-year-old. Please design it with an exciting development," "Please generate a gentle bedtime story for a 4-year-old. Please end with a calming ending," "Please generate a fun story for a 6-year-old with an animal protagonist," and "Please generate a short story about friendship for an 8-year-old." When adding word count or plot progression conditions to these prompts, the server can append descriptions, such as "Please keep the word count around 500 words" or "Please use a clear and easy-to-understand structure." Through this structured control at the prompt level, the server ensures that the generative AI model internally meets content length and structural constraints, thereby reducing the amount of post-processing corrections and improving overall computational efficiency.

[0225] In this implementation, the server does not rely on traditional template concatenation methods based on fixed rules. Instead, it uses a generative artificial intelligence model to probabilistically model text sequences, thereby learning complex language patterns in a high-dimensional parameter space. When invoking the model, the server dynamically adjusts the prompts based on user attributes, allowing different user groups to share the same set of model parameters while obtaining differentiated outputs. This approach avoids the redundant overhead of training or maintaining separate models for each user type. Since the generative artificial intelligence model has already learned multiple expressions and structural styles during the training phase, the server only needs to apply appropriate conditions to the prompts to trigger the corresponding generative patterns during the inference phase. This conditional control is achieved internally by directional changes in the vector space, rather than simply relying on manual if-else branches.

[0226] By breaking down generated content into segments and transmitting them in a streaming manner, the server refines the granularity of network communication and display processing to smaller data units. This finer granularity allows for more precise control over bandwidth usage and latency. For example, when network quality is poor, the server can shorten the length of each segment to reduce latency per transmission; when the network is smooth, the server can appropriately increase the segment length to reduce the number of communications. The server achieves adaptive control of communication load by dynamically adjusting the segmentation strategy. This mechanism is implemented by the computer analyzing network conditions and adjusting data block sizes at runtime, rather than simply manually presetting a fixed block size.

[0227] In this embodiment, the terminal not only performs simple display operations, but also adjusts interface parameters such as scrolling speed and pagination strategy based on the time interval and length information of received segments. For example, when the terminal detects that a segment arrives quickly, it sets the scrolling speed of the display area to be faster so that the user can follow the content in real time; when a segment arrives slowly, the terminal can insert a prompt message in the interface to remind the user that the content is still being generated. The terminal manages the receiving and display states through a local state machine, thereby achieving a smooth presentation of streaming data at the user interface level. This processing is not merely about automating business processes, but also about improving the technical performance of the human-computer interface at the levels of graphics rendering and interactive response.

[0228] In another embodiment of the invention, the server may not use a remotely provided generative artificial intelligence model, but instead deploy a pre-trained language generation model on local server hardware. In this embodiment, the server can utilize a graphics processing unit to accelerate matrix operations and store model weight parameters locally. With this local deployment, the server can further optimize the inference engine, for example, by reducing the computational load during model inference through methods such as quantizing weights, pruning redundant channels, and using low-bit-width arithmetic, thereby further improving response speed and reducing energy consumption.

[0229] In another alternative implementation, the server can incorporate a secondary decision-making module based on statistical features or machine learning into the content adaptation unit. In this module, the server calculates a series of quality indicators (such as vividness score, safety score, and age-appropriateness score) based on the parsed content elements and user attributes. Based on the comparison results of these indicators with thresholds, the server decides whether to regenerate part of the content or only perform partial replacement. This multi-level decision-making process employs algorithms such as decision trees or linear classifiers to achieve automatic screening and hierarchical control of the generated content quality, thereby improving the overall content reliability without increasing the burden of manual review by users.

[0230] In summary, this invention provides a specific computer implementation for generative artificial intelligence model invocation structure, prompt statement generation and encoding, streaming management and transmission of generated content, automatic adaptation processing based on user attributes, and terminal-side sequential display control through collaborative work between the server and terminal. The server optimizes resource utilization and communication load internally through specific data structures and processing flows, while the terminal enhances the interactive experience through interface control of streaming data. These specific technical means and system structure enable this invention to go beyond simple automation of manual editing work, simultaneously achieving comprehensive technical effects such as improved processing speed, enhanced adaptation accuracy, and optimized resource utilization across three computer technology levels: generative content processing, network transmission, and human-computer interaction.

[0231] use Figure 12 The processing flow is explained.

[0232] Step 1: Users select information categories and user attributes on the terminal.

[0233] Users operate the input module through the terminal's graphical interface and see the type and age lists on the screen or head-mounted display.

[0234] Input: User's click, swipe, and other operation events.

[0235] The terminal reads the currently selected information category (such as "adventure", "science fiction", "bedtime story" etc.) and the age value (such as "5 years old", "10 years old" etc.) in the user attributes based on the operation event, and saves it as structured data in local memory.

[0236] Output: A local data object containing information category fields and user attribute fields.

[0237] Step 2: The terminal sends data containing information categories and user attributes to the server.

[0238] The terminal takes the local data object from step 1 as input, calls the network communication module, and serializes the object into text-formatted request data.

[0239] Input: A data object containing information categories and user attributes.

[0240] The terminal processes the data object, encodes it into a request message, and sends it to the server's predetermined address using a network protocol via the wireless communication module.

[0241] Output: The request data stream transmitted through the communication network.

[0242] Step 3: The server receives and parses the request data from the terminal.

[0243] The server receives the request data stream sent in step 2 in the network interface module and passes it to the backend processing program.

[0244] Input: A request data stream containing information categories and user attributes.

[0245] The server performs data parsing operations on the requested data stream, extracting information category identifiers and user attribute values, and storing them in variables or data structures in the server's memory.

[0246] Output: Structured input parameters stored in the server's memory, including information categories and user attributes.

[0247] Step 4: The server generates a prompt statement based on the input parameters.

[0248] Based on the structured input parameters obtained in step 3, the server constructs natural language text in the prompt statement generation module.

[0249] Input: Information category (e.g., "Adventure") and age (e.g., 5) from the user attributes.

[0250] The server uses string concatenation and rule matching to convert age values ​​into age descriptions in the target language, maps information categories to topic descriptions in the target language, and then combines them into complete prompt text, for example: "Please generate an adventure story aimed at 5-year-olds." "Please generate a science fiction story aimed at a 10-year-old. Please design it to be as exciting as possible." "Please create a gentle bedtime story for a 4-year-old. End it in a way that helps them fall asleep peacefully." When the server needs to limit the length, it adds a word count condition to the prompt statement, such as "Please keep the word count to around 500 words". Data processing: The server transforms discrete category identifiers and numerical age parameters into semantically constrained information text based on predefined mapping tables and conditional rules.

[0251] Output: One or more prompt texts for generative artificial intelligence models.

[0252] Step 5: The server will input the prompt statement into the generative artificial intelligence model and start the generation process.

[0253] The server takes the prompt obtained in step 4 as input and calls the interface of the generative artificial intelligence model deployed on the computing resources.

[0254] Input: Prompt text.

[0255] In the model interface module, the server converts the prompts into a sequence of tags, maps each tag to a vector representation, and then feeds this vector sequence into a multi-layer transformer network structure. The server controls the model to perform forward propagation operations, calculating the attention matrix, context vector, and feedforward network output at each layer to obtain the conditional probability distribution of the next tag.

[0256] The server selects output markers from a probability distribution based on a set generation strategy (such as temperature parameters and candidate truncation) and gradually generates a text sequence.

[0257] Output: A stream of generated content tags, produced sequentially or in segments, in chronological order.

[0258] Step 6: The server organizes the tokenized stream output by the model into text segments.

[0259] The server takes the token stream obtained in step 5 as input, converts the token sequence into readable text in the text assembly module, and splits the complete text into multiple segments according to periods, newlines, or fixed-length rules.

[0260] Input: A stream of tags that will generate the content.

[0261] The server processes the data, decoding the marker sequence into a character sequence, and segments the string into multiple segments when a clause marker, paragraph marker, or a predetermined number of characters is detected. The server stores these segmented strings in a list structure corresponding to the session, and assigns a sequential number to each segment.

[0262] Output: Multiple segments of text information arranged in sequence.

[0263] Step 7: The server performs adaptation processing on the generated content.

[0264] The server takes the segment information list generated in step 6 as input and analyzes each segment in the content adaptation module in conjunction with user attributes.

[0265] Input: A list of segmented information text and user attributes (e.g., age).

[0266] The server performs string parsing and symbol sequence parsing on each segment of text, and counts features such as sentence length, vocabulary complexity, and total number of words. The server compares these features with preset thresholds to determine whether the expression difficulty and length are suitable for the corresponding age. If not, the server uses a replacement rule table to replace complex words with simple words, or splits long sentences into two or more short sentences, or removes redundant modifiers.

[0267] Data processing: Feature extraction, threshold comparison, and replacement and pruning based on mapping tables are performed on the text to obtain adapted segmented text.

[0268] Output: A list of adapted segmented information text that meets the user's attribute requirements.

[0269] Step 8: The server sends segmented information to the terminal in a streaming manner.

[0270] The server takes the adapted segmented information list output from step 7 as input and sends it segment by segment in sequence in the streaming module.

[0271] Input: A list of segmented information text after adaptation.

[0272] The server packages each segment into a transmission unit at the network layer and sends it to the terminal via the network interface device in list order. If necessary, a sequence number and end marker are added to each unit. During transmission, the server controls the transmission pace based on network conditions to achieve streaming output.

[0273] Output: Segmented information data streams transmitted over a communication network.

[0274] Step 9: The terminal receives and caches the segmented information sent by the server.

[0275] The terminal takes the data stream transmitted in the network in step 8 as input and receives it segment by segment in the local communication module.

[0276] Input: Segmented information data units arriving in sequence.

[0277] The terminal parses each data unit, extracts its segmented text and sequence number information, and appends and stores them in a local text buffer structure. The terminal maintains the segment order in the buffer structure to ensure the continuity of text during subsequent display.

[0278] Output: Segmentation information buffer data stored in the terminal's local memory.

[0279] Step 10: The terminal displays the generated content sequentially on the visual display device.

[0280] The terminal takes the segmented information buffer data obtained in step 9 as input and generates a text screen visible to the user in the display control module.

[0281] Input: Segmented information buffer data.

[0282] The terminal performs necessary formatting processing on newly received text segments, such as calculating line break positions based on screen size and font size, and appending the new text content to the end of the current display area. The terminal calls the graphics rendering interface to draw the updated text area onto the display module, allowing the user to see the continuously growing story content. If necessary, the terminal also displays progress indicators based on the number of segments or the reception progress.

[0283] Output: A dynamically updated text screen displayed on a visual display device.

[0284] Step 11: Users read on the terminal and make new requests as needed.

[0285] Users will read the story content displayed in step 10 and perceive the unfolding of the story through visual means.

[0286] Input: Text content displayed on the terminal screen or head-mounted display.

[0287] During the reading process, users can select actions such as "regenerate," "change category," or "adjust age" using buttons or gestures provided on the terminal, based on their personal experience. After detecting these actions, the terminal uses the new information category and user attributes as new input parameters and restarts the processing flow that began in step 1.

[0288] Output: New user selection information and a signal that triggers a new request to the server.

[0289] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.

[0290] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."

[0291] Existing story or picture book generation technologies based on generative artificial intelligence models typically generate one-time text content based solely on static conditions such as genre and age input by the user. Such systems suffer from the following technical problems: (1) The server side usually only performs simple parameter forwarding or template filling, lacking a structured prompt statement construction and dynamic control mechanism for generative artificial intelligence models. This makes it difficult to stably meet the fine-grained needs of users of different ages in terms of vocabulary difficulty, sentence length, content structure, etc., which limits the controllability and consistency of the system in terms of natural language generation quality.

[0292] (2) The server side lacks the ability to automatically and structurally process the generated results. It cannot programmatically split and reorganize long texts according to the page structure and scene division suitable for picture books or multi-page reading scenarios. It also cannot manage text content and prompts required for subsequent image generation or visual presentation in the same data structure, thus increasing the complexity of front-end and back-end collaboration and content reuse.

[0293] (3) Existing systems are weak in utilizing users’ emotional states. Servers usually do not have the ability to continuously feed back the emotional recognition results to the generative artificial intelligence model and automatically reconstruct elements such as scene composition, characters, expression intensity, and ending type of the generated story. As a result, the system cannot realize the dynamic evolution of the story content based on the user’s real-time emotional response, resulting in insufficient user experience and the inability to form a reusable emotion-content linkage control mechanism.

[0294] (4) Regarding visual presentation, the server lacks a unified mechanism to automatically generate visual prompts corresponding to the text content at the page level and store them as structured metadata associated with the work's structural data. This causes the text generation process to be separated from the subsequent image generation or interface presentation process, increasing the complexity of the overall system architecture and making it difficult to achieve integrated content generation and management on the server side.

[0295] Therefore, a new system architecture and program processing flow are needed to enable the server to: automatically generate highly semantically constrained prompts for generative artificial intelligence models, perform fine-grained processing of generated stories based on age and page structure, automatically generate page-level visual prompts and manage them in an integrated manner, and dynamically reconstruct and regenerate the work's structural data in conjunction with emotion recognition results. This will improve the control capabilities, scalability, and user interaction experience of natural language generation systems from a computer implementation perspective.

[0296] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.

[0297] In this invention, the server includes an acquisition and parsing device for receiving and parsing generation conditions, including the type of work and the target age, input by the user through an information processing device; a prompt construction and model calling device for automatically generating prompt statements for a generative artificial intelligence model based on the generation conditions and controlling the generative artificial intelligence model to generate story data; a structured editing device for editing the story data obtained from the generative artificial intelligence model according to the vocabulary difficulty, sentence length, expression content, and page division corresponding to the target age, and forming page-level work structure data; a visual prompt generation device for generating corresponding visual presentation prompt statements based on the text content of each page in the work structure data and associating the prompt statements as additional information for image generation or presentation with the work structure data; a recording and management device for recording and managing the generation conditions and generation history according to the user's save request or regeneration request; and an emotion linkage adjustment device for outputting the work structure data after the above editing and generation processing to the information processing device in a page-turning form, and combining it with a parsing device with emotion recognition function to analyze the user's emotional state, and automatically adjusting at least a part of the scene composition, characters, expression intensity, and ending type in the work structure data based on the emotional state. This allows for the formation of a complete control loop on the server side for generative artificial intelligence models. It can not only automatically construct high-quality prompts and generate structured story data that meets the needs of different age groups, but also integrate textual and visual prompts in a unified data structure. Furthermore, it can dynamically reconstruct and regenerate the structure of the work based on the user's emotional state, thereby significantly improving the controllability, adaptability, and interactivity of the natural language generation system at the computer implementation level, and enhancing the overall technical effect of content generation and presentation.

[0298] "Information processing device" refers to an electronic device that can communicate with a server and provide a user interface for receiving user input and displaying output content, including but not limited to terminal devices, computer devices, or other electronic devices with display and input functions.

[0299] "Generation conditions" refer to a set of parameter information input by the user and used by the server to control the generation process of a story or work. These parameters include at least the type of work and the target age, and may also include relevant parameters that control the output of the generative artificial intelligence model, such as protagonist settings, length preferences, and style requirements.

[0300] "Generative artificial intelligence models" refer to data processing models that are trained on large-scale data and can automatically generate natural language content or other content based on input text, including but not limited to deep learning models used for text generation.

[0301] "Prompt statements" refer to instructional texts automatically constructed by the server based on the generation conditions and provided to the generative artificial intelligence model. These texts are used to constrain or guide the generative artificial intelligence model to generate story data or other work data that meet predetermined requirements.

[0302] "Story data" refers to a collection of data, which is generated by generative artificial intelligence models based on prompts and describes plots, scenes, characters and their behaviors in natural language. It can be continuous text or text with structured tags.

[0303] "Work structure data" refers to structured data that the server generates after processing the story data. It is suitable for presentation by page or scene and includes at least page division information and the text content corresponding to each page. It may also include visual presentation information or metadata associated with that page.

[0304] "Display device" means a hardware or software component used to present an interface, text, images or other visual information to a user, including but not limited to a display screen, graphical user interface components or other visual output modules.

[0305] "Input device" refers to a hardware or software component used to receive user operations and input information, including but not limited to touch screens, keyboards, mice, microphones, or input controls in graphical interfaces.

[0306] "Control device" refers to the processing module in the server used to perform control logic such as parsing generation conditions, constructing prompt statements, and sending generation instructions to generative artificial intelligence models. It can be implemented by one or more processors and the software programs they execute.

[0307] The “editing device” refers to a module in the server used to process the story data output by the generative artificial intelligence model. This module adjusts the vocabulary difficulty, sentence length, expression content and page division according to rules such as the target age, thereby generating the work's structural data.

[0308] "Generation device" refers to a processing module in the server used to generate visual prompts based on the page text content of the work's structural data, and associate the prompts with the corresponding page. It can be used to support subsequent image generation or interface presentation.

[0309] "Recording device" refers to a storage and management module in a server used to store and manage generation conditions, generation history, and related work structure data. It can be implemented using a database or other storage media to support data querying, reuse, and version management.

[0310] "Output device" refers to a module in a server used to send the work's structural data and its associated prompts to an information processing device via a predetermined communication protocol and drive the device to display it, including a network communication interface and its control program.

[0311] "Analysis device" refers to a processing module with emotion recognition function, which is used to receive input information related to the user (including but not limited to text feedback, physiological signals or interactive behavior) and identify the user's emotional state based on the information.

[0312] "Emotional state" refers to the classification or quantitative representation of a user's psychological or emotional characteristics at a specific moment, inferred by the analysis device based on relevant user information. This includes, but is not limited to, emotional types such as happiness, tension, fear, and boredom, or corresponding numerical indicators.

[0313] "Adjustment device" refers to a processing module in the server used to modify, replace, or rearrange elements such as scene composition, characters, intensity of expression, and type of ending in the work's structural data based on the user's emotional state identified by the analysis device, thereby realizing dynamic adjustment of the story content.

[0314] "Scene composition" refers to the combined structure in the work's structural data used to describe the story's environment, time, location, and related background elements, including specific scene settings and environmental descriptions under page divisions.

[0315] "Entities appearing in the story" refers to the collection of characters or things that appear in the story data or work structure data, including protagonists, supporting characters and other objects with narrative functions. Their attributes may include names, personality traits, appearance characteristics, etc.

[0316] "Expressive intensity" refers to parameters or features in story data that reflect the emotional tension or atmosphere of the narrative, including the degree of conflict, tension, and humor, and is related to word choice, sentence structure, and plot arrangement.

[0317] "Ending type" refers to the category of the way the story ends in the work's structural data, including but not limited to happy ending, open ending, educational ending, etc., which are used to describe the overall direction of the story's final scene.

[0318] "Visual presentation prompts" refer to descriptive text generated by the server for each page or scene based on the corresponding text content. These texts are used to guide the image generation model or interface design in generating corresponding visual content and specify visual elements such as the main subject, environment, color tone, and style of the image.

[0319] "Additional information for image generation or presentation" refers to auxiliary data associated with the work's structural data and used by the image generation model or front-end presentation module. This includes metadata used to control visual output, such as visual presentation prompts, style parameters, and color preferences.

[0320] "Generation history" refers to the record of each generation process related to a specific user or a specific work, including information such as the generation conditions used, generation time, version of the generative artificial intelligence model, and summary of the main output content, which is used to support backtracking, comparison, and regeneration control.

[0321] The embodiments of the present invention will be described in conjunction with the information processing system structure described in Appendix 1 to 3. In the following description, the subject is limited to one of "server," "terminal," or "user," and the system of the present invention will be specifically explained from the aspects of hardware composition, software modules, data structure, algorithm flow, and causal relationship of technical effects.

[0322] System Overall Composition Servers can be composed of general-purpose computing devices, such as rack-mounted computers or cloud computing nodes equipped with multi-core central processing units (CPUs), graphics processing units (GPUs), main memory, and non-volatile memory. Servers can run Linux-based operating systems, such as Ubuntu. Servers can communicate with multiple terminals via a local area network (LAN) or a wide area network (WAN).

[0323] The terminal can be a tablet computing device, a smartphone, or a laptop computing device, etc. The terminal can run a mobile operating system or a desktop operating system, such as Android or a universal desktop operating system. The terminal may include a touchscreen display, a speaker, a microphone, and a wireless communication module.

[0324] Users can interact with the server through an application running on the terminal. This application can be a native application (such as one implemented on Android using Kotlin or Java), a cross-platform application (such as one implemented using Flutter or React Native), or a web application accessed through a web browser.

[0325] The server can use backend frameworks at the application layer, such as Python's FastAPI or Django framework, Java's Spring Boot framework, or JavaScript's Node.js and Express framework. The server can use relational database management systems (such as MySQL or PostgreSQL) and key-value caching systems (such as Redis) to store and manage project structure data, generation conditions, and generation history.

[0326] Server-side functional modules A server can logically include the following functional modules, which can be implemented by software programs executed by one or more processors: 1. Generation Condition Acquisition and Parsing Module The server can receive generation conditions from the terminal, including the type of work and the target age. The server can use the request object provided by the web framework to obtain the HTTP request body, and then parse the request body using a JSON parsing library to obtain a structured generation condition object. After parsing, the server can perform validity checks, such as determining whether the target age is within a preset range and whether the type of work is in the supported list.

[0327] 2. Prompt Statement Construction and Model Calling Module The server can construct prompts based on the generated conditions. The server can maintain a template library, where each template corresponds to a different age group and a different type of content. For example, the server can select the following template for adventure content for a 3-year-old user and replace the placeholders with specific parameters: "You are an author who is good at creating picture book stories for 3-year-old children."

[0328] Please create a Chinese story according to the following requirements: 1. Story type: Adventure.

[0329] 2. Main character: A cute little animal named Little Rabbit.

[0330] 3. Target audience: 3-year-old children.

[0331] 4. Language requirements: Use very simple vocabulary and short sentences; avoid using idioms and complex words.

[0332] 5. Plot requirements: The story should be lighthearted and fun, and should not contain any bloody or terrifying scenes.

[0333] 6. Length: Keep it under 500 words.

[0334] 7. Ending: It must be a happy and heartwarming ending.

[0335] Please output the complete story text directly. The server can send the prompt as text input to the generative AI model. The server can invoke a large-scale neural network model deployed on a remote inference service. The generative AI model can be an autoregressive language model based on the Transformer architecture, which can include multiple layers of self-attention sublayers and feedforward sublayers, each of which can include a multi-head attention mechanism.

[0336] The server can specify model parameters at the time of invocation, such as maximum output length, sampling temperature, and top-k or top-p sampling thresholds. The server can communicate with the model inference service via HTTPS, transmitting prompts to the model and receiving the story data output by the model.

[0337] Structure and Training of Generative Artificial Intelligence Models The server can use a pre-trained and fine-tunable generative artificial intelligence model. This model can have the following structural features and training methods: 1. Model Structure Generative artificial intelligence models can be multi-layered Transformer language models, each layer including: - Multi-head self-attention unit, used to calculate the correlation between positions in the input sequence; - A feedforward fully connected network is used to perform non-linear transformations on the latent vector at each location; - Residual connections and layer normalization units are used to improve gradient propagation stability.

[0338] The model can use word segmentation methods such as byte-pair encoding to encode Chinese text into a discrete token sequence. The model can use an embedding matrix to map the tokens to vectors and use positional encoding to represent the sequence's positional information.

[0339] 2. Training Data and Objectives Generative AI models can be pre-trained unsupervised on large-scale, multi-domain Chinese corpora. The training objective can be to minimize the negative log-likelihood loss of the autoregressive language model, i.e., to minimize the prediction error for the next label at each time step.

[0340] 3. Fine-tuning and alignment The server can fine-tune the model based on a corpus of children's stories to enhance its performance in the children's story domain. During fine-tuning, the server can use a specific dataset (e.g., age-annotated story text) and employ cross-entropy loss to update the model parameters using gradient descent. The server can also use reinforcement learning methods based on human feedback after fine-tuning to align the model output for quality.

[0341] 4. Specific control during the reasoning stage During the inference phase, the server can control the diversity and stability of the output by setting the sampling temperature and top-p parameters. The server can also add explicit structural guidance information to the prompts, such as requiring segmented or scene-specific descriptions, for subsequent structured processing.

[0342] By using this combination of structured hints and controlled sampling, the server can obtain more controllable story data with fewer syntax errors and more consistent style without significantly increasing computational complexity, thereby improving generation accuracy and reducing post-processing burden.

[0343] Structured editing and age-adaptation of story data The server can perform structured editing on the story data output by the model to generate work structure data.

[0344] 1. Text cleaning and sentence segmentation The server can use sentence segmentation algorithms to process the story data. For example, the server can segment long texts into sentences based on punctuation detection (periods, question marks, exclamation marks, etc.) and Chinese grammar rules. The server can use string manipulation and regular expressions to remove irrelevant phrases (such as "Okay, here is the story") and extra blank lines from the model output.

[0345] 2. Age-appropriate vocabulary replacement The server can maintain a complex vocabulary dictionary and a corresponding simple vocabulary mapping table. For example, "explore the unknown territory" can be mapped to "go see a new place," and "fear" can be mapped to "be afraid." The server can use word segmentation tools (such as a general Chinese word segmentation library) to segment the story data and scan each word to see if it is in the complex vocabulary. If it is, it replaces the word with a preset simple vocabulary and adjusts the syntax according to the context.

[0346] 3. Sentence length control The server can set maximum sentence length thresholds for different age groups, such as setting each sentence to no more than 30 characters for 3-year-olds. After segmenting sentences, the server can calculate sentence length and further break down sentences exceeding the threshold according to pauses, commas, or conjunctions to form multiple shorter sentences. This can significantly improve comprehension for younger users.

[0347] 4. Page partitioning strategy The server can divide the story into multiple pages based on the total word count, number of sentences, and scene change cues. The server can use a greedy algorithm to sequentially add sentences to the page, checking for paragraph breaks or scene transitions as page breaks occur when the page approaches its expected word count. The server can record fields such as `page_index` and `page_text` in the data structure.

[0348] Through the aforementioned structured editing, the server not only generates text more suitable for the target age group, but also internally forms work structure data with clear page boundaries, thereby improving the efficiency of subsequent rendering and retrieval.

[0349] Generation and association of visual presentation prompts The server can automatically generate visual prompts for each page of the work's structural data and perform structured association.

[0350] 1. Scene Feature Extraction The server can perform simple semantic analysis on the text content of each page to extract elements such as main characters, environment, and emotions. The server can use a combination of rules and statistics, for example: - Use entity recognition tools to extract character names (e.g., "Little Rabbit"); - Identify environment types (such as "forest", "seaside", "room") based on a keyword list; - Estimate scene emotions (such as "happy" or "nervous") based on an emotion dictionary.

[0351] 2. Construction of visual cues The server can use the above elements to apply a template to generate prompts. For example, for the scene "a little rabbit in front of a little house by the forest," the server can generate: "A cute little rabbit stands at the door of a small house, next to a green forest. The sky is bright with a few small white clouds, their colors bright and warm." The server can store the visual prompt statement as text in the image_prompt field of the corresponding page in the artwork structure data.

[0352] 3. Related to image generation technologies When the system works in conjunction with an image generation service, the server can send visual cues as input to the image generation model (e.g., an image generation network based on a diffusion model) to automatically generate illustrations. Alternatively, the server can simply provide the cues to the terminal or a human illustration system for drawing reference.

[0353] By managing text and visual cues uniformly on the server side, the system can achieve coupled control of the text and image generation process, thereby reducing the probability of rendering errors caused by data inconsistency and improving the overall maintainability of the system.

[0354] Emotion recognition and dynamic adjustment of work structure The server can use an emotion recognition module to identify the user's emotional state and dynamically adjust the structure data of the work based on the recognition results.

[0355] 1. Sentiment Data Acquisition The terminal can send user interaction data during the reading process to the server, such as page dwell time, rapid page turning behavior, and user manual feedback (e.g., button selection for "too scary" or "too boring"). The terminal can also send voice data or emoticons. In one embodiment of the invention, the server can use only the interaction behavior data and is not required to rely on acoustic or visual data.

[0356] 2. Emotion Recognition Algorithm The server can use a classification model to identify sentiment states. This classification model can be a shallow neural network or a gradient boosting tree-based model. Input features can include: - Statistical characteristics of time spent on each page; - Is there frequent, rapid page-turning behavior? - Explicit rating tags for specific pages by users; - Key events in the conversation history (such as triggering the "Too scary" button multiple times).

[0357] The server can standardize these features and input them into the classification model. The server can train the model using cross-entropy loss and fine-tune the hyperparameters using a validation set. The model output can be a probability distribution of multiple emotion categories, such as "happy," "nervous," "fearful," and "bored."

[0358] 3. Rules for Adjusting the Structure of the Work The server can apply a set of rules to adjust the structure data of the work based on the emotion recognition results: - When a user is identified as "fearful", the server can reduce the proportion of tense scenes in subsequent pages, for example, by replacing dangerous scene descriptions with mild scene descriptions, or by introducing safety roles in advance; - When a user is deemed "bored", the server can insert or enhance plot conflict, such as adding new adventure objectives or humorous events; When a user is judged to be "happy", the server can maintain the current narrative pace and only make minor adjustments to the dialogue content in certain areas.

[0359] The server can maintain an internal library of replacement clips, each clip in which a predefined emotion tag and applicable age range can be defined. The server can select an appropriate replacement clip based on the emotion tag and perform text replacement or addition / deletion operations on the corresponding page.

[0360] In another implementation, the server can reconstruct new prompts based on emotion recognition results, issue supplementary generation instructions to the generative AI model to generate new text paragraphs that better match the current emotional state, and then insert the new text paragraphs into the work's structural data. This allows for on-demand local regeneration without having to recalculate all pages, thereby reducing computational load and waiting time.

[0361] The technical effects of data structures and data management The server can store the work's structure data, generation conditions, and generation history in a relational database. The server can be designed with standardized table structures, for example: - Works Table: Stores work identifier, user identifier, creation time, and overall attributes; - Page table: Stores the text content, page index, and corresponding visual cues for each page; - Condition table: Stores the generation conditions associated with the work (work type, age, etc.); - History table: Stores the time point of each regeneration and parameter changes.

[0362] Servers can improve query speed through index optimization and preloading techniques. Since the work's structured data is already structured by page and field, the server can send only the currently viewed page and a few adjacent pages to the terminal, thereby reducing network traffic and communication load.

[0363] The server employs a structured data management approach, which differs from solutions that simply transmit long texts, and can achieve the following technical benefits: - Reduce the size of a single network response, improve transmission efficiency, and reduce latency; - Facilitates updating and cache hits on local pages, thereby improving overall system throughput; - Facilitates the execution of batch analysis and automated testing on the server side, thereby improving system reliability.

[0364] Terminal-side presentation and interaction The terminal can receive the work structure data sent by the server and present it graphically according to the page structure. The terminal can use page container components (such as ViewPager on Android and pagination components in cross-platform frameworks) to implement horizontal swiping page turning.

[0365] The terminal can cache portions of the current work locally, reducing the number of times the same data is repeatedly requested from the server. The terminal can send user modifications or annotations to a page as differential data to the server, and the server can store the original version and the user-edited version separately in the recording module.

[0366] The terminal can send a regeneration request to the server, which can include the user's selection of a new genre or target age. Upon receiving the regeneration request, the server can reconstruct the prompts based on the original and new generation conditions, invoke the generative artificial intelligence model, and partially rewrite the work's structural data.

[0367] Explanation of technical effects and causal relationships Through the specific data structure design, prompt statement construction method, model invocation strategy, emotion linkage adjustment logic, and structured storage scheme described above, the server can achieve technical effects that differ from simple manual task automation in the following aspects: 1. Improved accuracy and controllability By generating semantically constrained prompts based on age and genre, combined with age-appropriate vocabulary replacement and sentence length control, the server can significantly reduce the probability of generative AI models producing age-inappropriate content or grammar that doesn't conform to the target audience's reading habits. Consequently, the quality and controllability of story data are improved.

[0368] 2. Improved processing speed and computational efficiency Through page-level structured editing and local regeneration mechanisms, the server can regenerate or replace only the affected pages without regenerating the entire work, thereby reducing the number of calls to generative AI models and the overall computational load.

[0369] 3. Improved data management and communication efficiency By breaking down story data into work structure data and managing text and visual cues in a unified manner, the server can send only the necessary fields to the terminal for the pages that need to be displayed or updated, thereby reducing network bandwidth consumption and latency, and facilitating the application of caching mechanisms.

[0370] 4. Emotionally driven dynamic adjustment ability Through emotion recognition models and rule bases, the server can dynamically adjust the story content during user reading. This mechanism of controlling the story structure based on the user's emotional state is difficult to achieve in traditional systems based on static rules. This mechanism enables the system to achieve content adaptive control based on real-time feedback, thus improving the system's intelligence level.

[0371] 5. Interpretability and controllability of the model's internal processing By explicitly defining technical details such as model training objectives, loss functions, and sampling strategies, the server transforms model invocation and output control from a "black box operation" into an engineering process that allows for fine-tuning of parameters and structural optimization for different application scenarios. This differs from the traditional approach of treating models merely as uninterpretable components, and facilitates continuous optimization of the system in engineering practice.

[0372] Multiple implementation forms and variations In the first implementation, the server can simply call a generative artificial intelligence model deployed on an external cloud. All model calculations are performed on external resources, and the server is responsible for generating prompts, post-processing, and managing data.

[0373] In the second implementation, the server can deploy a scaled-down version of the language model locally for offline or low-latency scenarios. The server can switch between the local model and the cloud model based on load and latency requirements, or adopt a two-tier model architecture, where the local model is used for coarse-grained generation and the cloud model is used for high-quality re-polishing.

[0374] In the third implementation, the server can be configured with different template libraries and vocabulary mapping tables for different languages ​​or regions to adapt to multilingual and multicultural environments, while using the same data structure and module interface to ensure the consistency of the system architecture.

[0375] In the fourth implementation, the server can select visual cue generation strategies of different complexities based on the performance of the terminal device. For example, it can generate more detailed visual cue statements for high-performance terminals and generate simplified descriptions for low-performance terminals, in order to balance computational overhead and image generation quality.

[0376] Through the above-mentioned various implementation forms and modifications, the system of the present invention can achieve stable operation under different hardware conditions and application scenarios, while maintaining its technical advantages in generation control, structured management and emotional linkage.

[0377] use Figure 13 The processing flow is explained.

[0378] Step 1: The user enters the generation conditions on the terminal.

[0379] Users can select or enter the type of work (e.g., "adventure" or "bedtime") and target age (e.g., "3 years old" or "5 years old") in the application interface on the terminal, and optionally enter additional information such as the protagonist's name and length preference.

[0380] Input: Raw input events generated by user actions in interface controls (drop-down lists, text boxes, radio buttons, etc.).

[0381] The terminal collects and organizes these raw events to generate an internal parameter object, which includes fields such as "genre", "age", "main_character", and "length_preference".

[0382] Output: Structured generated conditional data on the terminal side, used for subsequent network transmission.

[0383] Step 2: The terminal performs local verification and serialization of the generation conditions.

[0384] The terminal checks whether the generation conditions include the required fields (artwork type and target age) and verifies whether the age value is within the preset range and whether the art type is in the supported list. If a missing or invalid value is found, the terminal displays an error message on the interface and prevents further processing.

[0385] Input: The structured generation condition data generated in step 1.

[0386] After the terminal passes the verification, it uses a JSON serialization library to encode the generated condition object into a JSON string and generates an HTTP request body locally.

[0387] Output: Valid JSON request data to be sent to the server.

[0388] Step 3: The terminal sends a generation request to the server.

[0389] The terminal uses a network communication module (e.g., using HTTPS protocol and built-in HTTP client library) to encapsulate the JSON request data obtained in step 2 in an HTTP POST request and sends it to the predefined interface address of the server. After sending the request, the terminal starts a loading animation and waits for the server's response.

[0390] Input: Generate conditional request data in JSON format.

[0391] The terminal encapsulates the data at the transport layer and sends it out, generating a network data packet.

[0392] Output: The network request message sent to the server, and the terminal entering the "waiting for server response" state.

[0393] Step 4: The server receives the request and parses the generation conditions.

[0394] The server obtains the HTTP request body through the request object provided by the backend framework, parses the request body using a JSON parsing library, and extracts parameters such as the genre of the work, target age, and main character's name. After parsing, the server performs a validity check, and can return an error response if missing fields or invalid values ​​are detected.

[0395] Input: An HTTP request message from the terminal and its JSON request data.

[0396] The server parses and processes the JSON string, converting the text into a key-value pair structure and constructing an internal conditional object.

[0397] Output: The standardized generated condition object inside the server, which serves as the input for constructing subsequent prompt statements.

[0398] Step 5: The server constructs a prompt statement based on the generated conditions.

[0399] Based on the genre and target age, the server selects a suitable prompt template from the template library and fills in the specific parameters from the generation conditions into the template placeholder positions. For example, when the genre is "adventure," the target age is "3 years old," and the protagonist is a "little rabbit," the server constructs the following prompt statement: "You are an author who is good at creating picture book stories for 3-year-old children."

[0400] Please create a Chinese story according to the following requirements: 1. Story type: Adventure.

[0401] 2. Main character: A cute little animal named Little Rabbit.

[0402] 3. Target audience: 3-year-old children.

[0403] 4. Language requirements: Use very simple vocabulary and short sentences; avoid using idioms and complex words.

[0404] 5. Plot requirements: The story should be lighthearted and fun, and should not contain any bloody or terrifying scenes.

[0405] 6. Length: Keep it under 500 words.

[0406] 7. Ending: It must be a happy and heartwarming ending.

[0407] Please output the complete story text directly. Input: The generation condition object obtained in step 4 (including the type of work, target age, protagonist's name, etc.).

[0408] The server performs a string template replacement operation, inserting parameter values ​​into a predefined template to generate a single, continuous Chinese instruction text.

[0409] Output: The complete text of the prompts for generative artificial intelligence models.

[0410] Step 6: The server calls a generative artificial intelligence model to generate story data.

[0411] The server takes the prompt generated in step 5 as input and sends it to the generative AI model via a model call interface. This model can be an autoregressive language model based on the Transformer architecture. The server sets inference parameters (such as maximum output length, sampling temperature, top-p value, etc.) and calls a remote inference service via HTTPS. The generative AI model performs word embedding, attention calculation, and probability prediction on the prompt, gradually outputting the story text.

[0412] Input: The prompt text constructed by the server, and the inference configuration parameters.

[0413] The server encapsulates the prompt statement into a model input structure. The generative AI model performs matrix multiplication, nonlinear transformation, and probability sampling calculations on this structure, and outputs a story text string.

[0414] Output: The original story data text returned by the generative artificial intelligence model, stored in the server's memory.

[0415] Step 7: The server performs text cleaning and sentence segmentation on the story data.

[0416] The server performs a cleaning operation on the raw story data returned in step 6, removing irrelevant prefixes such as "Okay, here is the story," and deleting redundant blank lines and abnormal characters. Subsequently, the server segments the long text into several sentences based on punctuation and preset rules, forming a list of sentences.

[0417] Input: A long text string containing the model's original output.

[0418] The server uses regular expressions and string splitting algorithms to scan, replace, and segment the text, generating an array of sentences.

[0419] Output: A cleaned and segmented sequence of story sentences, providing a foundation for subsequent age-appropriate matching and page segmentation.

[0420] Step 8: The server performs word replacement and sentence length control based on the target age.

[0421] The server performs age-adaptation processing on the sentence sequence from step 7. First, the server uses a Chinese word segmentation tool to segment each sentence, transforming it into a word sequence. Then, it checks if each word is in a complex vocabulary. If a word is marked as difficult, the server replaces it with a corresponding simple word based on a mapping table. The server also calculates the character length of each sentence. When the length exceeds a threshold set for the target age, the server splits the long sentence into two or more sentences based on natural pauses such as commas and conjunctions.

[0422] Input: The sequence of story sentences after sentence segmentation, the target age parameter, a complex vocabulary, and a vocabulary mapping table.

[0423] The server performs word segmentation, dictionary lookup, and sentence recombination operations, and outputs a sentence sequence that meets the language difficulty requirements for the target age group.

[0424] Output: A sequence of age-appropriate story sentences with simplified vocabulary and controlled sentence length.

[0425] Step 9: The server divides age-appropriate story sentences into pages and generates work structure data.

[0426] The server uses a greedy strategy to distribute the sentence sequence obtained in step 8 to multiple pages based on a preset maximum number of characters or sentences per page. When approaching the maximum number of characters or sentences per page, the server prioritizes sentence boundaries or natural paragraph boundaries as page break points. The server assigns a page index to each page and merges the sentences on that page into one or more paragraphs.

[0427] Input: A sequence of story sentences that have been age-appropriated and page segmentation parameters (maximum number of characters per page, maximum number of sentences, etc.).

[0428] The server performs cumulative counting and boundary judgment operations, maps sentences to different page structures, and constructs a list of pages containing page_index and page_text.

[0429] Output: Represents the page-level work structure data for the entire work, with each item containing the page index and page text content.

[0430] Step 10: The server generates visual prompts for each page.

[0431] The server performs simple semantic analysis on the text of each page in step 9, extracting key characters, environment, and emotional information, such as identifying elements like "forest" and "rabbit" using a keyword list. Based on the extracted elements, the server applies a visual description template to generate descriptive text that can be used for image generation or illustration reference, for example: "A cute little rabbit stands at the door of a small house, next to a green forest. The sky is bright with a few small white clouds, their colors bright and warm." Input: The text content of each page, as well as the character dictionary, environment dictionary, and emotion dictionary.

[0432] The server performs keyword matching and template filling operations to construct corresponding visual description text for each page.

[0433] Output: Visual prompts corresponding to each page, which are appended to the artwork structure data as the image_prompt field.

[0434] Step 11: The server stores generation conditions, generation history, and work structure data.

[0435] The server writes the generation conditions, generation time, user identifier (if any), and the work structure data obtained in steps 9-10 into the database. The server records the work identifier and basic attributes in the work table, the page text and visual prompts in the page table, and the parameters and time of this generation request in the history table.

[0436] Input: The generation conditions for this request, the structure data of the work (including page text and visual cues), and meta-information such as user identification.

[0437] The server performs database insert or update operations, converts the data into table records, and persists them to storage media.

[0438] Output: Persistent data with unique work identifiers and historical records, facilitating subsequent querying and regeneration.

[0439] Step 12: The server generates a response based on the work's structure data and sends it to the terminal.

[0440] The server reads the generated artwork structure data from the database or memory, constructs a response object containing the artwork title, target age, artwork category, page content, and corresponding visual cues. The server serializes this response object into a JSON string, encapsulates it in an HTTP response message, and returns it to the client over the network.

[0441] Input: Processed and stored work structure data and work metadata.

[0442] The server performs data selection, object construction, and JSON serialization operations to generate response data that can be transmitted over the network.

[0443] Output: An HTTP response message containing page-level content and visual cues.

[0444] Step 13: The terminal parses the server response and renders the page.

[0445] The terminal receives the HTTP response from the server, uses a JSON parsing library to parse the response body into an internal data structure, and extracts the work title, target age, and page list from it. The terminal creates a corresponding UI component for each page, draws the page text onto text controls, and displays images generated based on `image_prompt` or preset images when needed. The terminal is configured with page-turning controls, allowing users to turn pages by swiping or using buttons.

[0446] Input: JSON format response data sent by the server.

[0447] The terminal performs parsing and interface layout operations, mapping structured data to interface elements and generating an interactive picture book reading interface.

[0448] Output: The page-turning picture book content displayed on the terminal screen, along with a list of page objects in memory associated with it.

[0449] Step 14: Users read and generate interactive data on the terminal.

[0450] Users read picture books page by page on the device by swiping the screen and clicking page-turning buttons. Users can click feedback buttons such as "like," "terrible," and "boring" on certain pages, or submit comments through text input boxes. The device tracks the time spent on each page and the page-turning speed in the background, and then aggregates this behavioral data with user feedback.

[0451] Input: User's interaction events and reading behaviors on the terminal interface (touch events, button clicks, dwell time, etc.).

[0452] The terminal performs event listening and time statistics calculations, converting raw events into structured interactive data records.

[0453] Output: Page-level interaction data and emotion-related feedback data, ready to be sent to the server for emotion recognition.

[0454] Step 15: The server receives interactive data and performs emotion recognition.

[0455] The server receives the interaction data uploaded in step 14 from the terminal, parses it, and extracts features, such as average dwell time per page, number of quick page turns, and click frequency of different emotion feedback buttons. The server inputs these features into an emotion recognition model, which can be a pre-trained classifier (e.g., a shallow neural network or a gradient boosting tree). The server runs the model to calculate the user's emotion state category and its probability distribution in the current session.

[0456] Input: Interaction data records uploaded by the terminal (dwell time, page turning behavior, button feedback, etc.).

[0457] The server performs feature engineering operations (normalization, statistical aggregation) and model inference operations (vector multiplication, activation functions, decision tree traversal, etc.) to output a sentiment label and its corresponding probability.

[0458] Output: Represents the classification results of the user's emotional state, such as "happy", "fearful", "bored", etc., and the corresponding confidence scores.

[0459] Step 16: The server dynamically adjusts the structure data of the work based on the emotional state.

[0460] Based on the emotional state obtained in step 15, the server searches for predefined adjustment rules. For example, when the emotional state is "fear" and the current work is for young users, the server selects a milder scene description from the replacement fragment library to replace the tense plot in the current or subsequent pages; when the emotional state is "bored," the server inserts a fragment with more action or humor. For cases requiring text regeneration, the server can construct local prompts, call a generative AI model to generate new fragments only for specific pages or scenes, and then insert those fragments into the work's structure data.

[0461] Input: Current work structure data, user emotional state classification results, replacement segment library, and adjustment rule set.

[0462] The server executes matching rules, selects fragments or locally generated operations, and replaces, inserts, or deletes text on specified pages to generate an updated version of the work's structure data.

[0463] Output: New work structure data, adjusted according to the user's emotional state, for resending to the terminal.

[0464] Step 17: The terminal receives the updated work structure data and refreshes the display.

[0465] The terminal receives the updated work structure data sent by the server after processing in step 16, and parses out the page indexes and corresponding content that need to be updated. Without reloading the entire work, the terminal only replaces or re-renders the content of the affected page UI components, so that the user sees the story content optimized according to their emotional state when they continue reading.

[0466] Input: Updated work structure data pushed or returned by the server (may only include the changed pages).

[0467] The terminal performs differential calculations (comparing the old and new page indices) to update and redraw the text and images of the corresponding interface components.

[0468] Output: The dynamically adjusted picture book content presented on the terminal, thereby achieving a continuous reading experience based on emotion.

[0469] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0470] With the widespread application of generative artificial intelligence models in the field of content generation, existing computer-based story generation systems, although capable of automatically generating text based on user-input topics or keywords, still have significant technical limitations in the following aspects, making it difficult to fully utilize the processing power of computer systems.

[0471] First, existing systems mostly rely solely on static genre and age parameters to invoke generation models, lacking mechanisms for acquiring and utilizing users' real-time emotional states. This results in generated content that doesn't match the user's current mood and an inability to automatically adjust the generation logic based on the user's emotional feedback. This architecture, lacking emotional loop control, prevents the computer system from adaptively optimizing the content generation strategy during operation.

[0472] Second, existing systems often use fixed templates or simple splicing to construct prompts, failing to fully utilize story elements in the database and random selection mechanisms. This results in insufficient structure, diversity, and consistency with user conditions in the prompts input to the generative AI model, thus affecting the convergence effect and output quality of the generative model, and also leading to low efficiency in the utilization of computing resources.

[0473] Third, existing systems typically output the generated results directly as a linear text, with almost no process for re-analyzing or controlling the sentiment of the generated content. Furthermore, they lack a closed-loop control mechanism to reconstruct and rewrite the content using prompts to drive the generative AI model for secondary or re-generation. Therefore, computer systems cannot perform programmatic quality evaluation and automatic correction after the model's black-box output.

[0474] Fourth, when converting generated text into ebooks or picture books, existing systems generally use manual layout design or simple segmentation, lacking a structured book data generation process based on page granularity, as well as a combined processing flow that automatically generates image prompts based on the content of each page and calls the image generation model. This results in loose coupling between the text generation module and the image generation module at the system architecture level, which is not conducive to achieving a unified data flow and automated page-level layout control within the computer.

[0475] The aforementioned problems prevent traditional story generation systems from fully integrating the unified processing chain of "user input—sentiment analysis—prompt generation—multimodal content generation—sentiment feedback re-optimization—structured book generation" in terms of computer architecture and software workflow. This limits the improvement of interactivity, personalization, and system performance in content generation based on generative artificial intelligence models. Therefore, it is necessary to propose a new computer implementation scheme that improves the autonomy, scalability, and processing efficiency of the content generation system at the computer technology level by refining the server-side processing flow, prompt generation strategy, sentiment analysis and regeneration control logic, and book structure data generation method.

[0476] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.

[0477] In this invention, the server includes: means for acquiring user operation information and generating a user interface for the user to specify information categories and target age on a terminal device; means for acquiring story element candidates from a storage device according to the information categories and the target age and selecting at least one story element through random processing; means for automatically generating prompt statements for inputting into a generative artificial intelligence model based on conditional information including the information categories, the target age, and the story elements; means for inputting the prompt statements into the generative artificial intelligence model to generate story content and obtaining the generation result; means for performing emotion analysis processing based on biometric information obtained by an imaging device and a voice acquisition device to infer the user's emotional state and modifying or regenerating at least a portion of the story content according to the emotional state; means for generating rewriting prompt statements containing existing story content and driving the generative artificial intelligence model to regenerate again when the emotion analysis processing result does not meet a predetermined target emotional condition; and means for generating book structure data including multiple page information and visual information layout information corresponding to each page for the modified or regenerated story content, and generating image generation prompt statements based on the description content of each page to call the image generation artificial intelligence model to obtain image data when needed. This allows for the formation of a closed-loop data processing chain within the server, encompassing user condition acquisition, random selection of story elements, construction of structured prompts, initial generation and regeneration based on emotional feedback, culminating in automatic layout of page-level book structure and multimodal content. This enables precise control and adaptive optimization of the generative artificial intelligence model invocation process at the computer technology level, improving emotional matching accuracy, diversity, and the degree of automation in layout generation, thereby enhancing the overall system's processing efficiency and user experience.

[0478] "System" refers to an entire system consisting of one or more information processing devices, terminal devices, and communication connections between them, used to perform the data processing and content generation functions described in this invention.

[0479] "User action information" refers to data generated when users perform actions such as selection, clicking, and text input through input devices or graphical user interfaces, which is used to represent user intentions or preferences. This includes information classification, target age, story element selection, and evaluation feedback.

[0480] "Information classification" refers to the parameters used to categorize the content to be generated, representing the general theme or attributes of the content, such as adventure, fantasy, science fiction, etc.

[0481] "Target age" refers to a parameter used to characterize the age group of the intended readers of the generated content. It is age-related information used to control the difficulty of the content, vocabulary selection, and plot complexity.

[0482] "User interface" refers to the graphical interface or interactive screen presented on a terminal device, which is a display and interactive structure for users to input information, select options, and browse results.

[0483] "Terminal device" refers to an electronic device that is directly operated by the user, used to display the user interface, collect user operation information and biometric information, and communicate data with the server, such as a mobile terminal, tablet terminal or other display terminal.

[0484] "Storage device" refers to data storage hardware or its logical combination used to store information such as candidate story elements, generated result data, and book structure data, and may include database systems, local storage or external storage resources.

[0485] "Story element candidates" refers to the set of basic building blocks used to construct story content, which are pre-stored in a storage device, including candidate data such as character settings, scene settings, plot templates, and theme information.

[0486] "Story elements" refer to one or more specific elements selected from the story element candidates, used to specify the characters, scenes, plots, and other content that constitute the story in prompt statements or generation logic.

[0487] "Random processing" refers to a processing method that selects one or more elements from a candidate set in an unordered manner, based on probability or pseudo-random algorithms, rather than following a fixed order or fixed rules.

[0488] "Conditional information" refers to a set of input parameter data used to constrain and guide generative artificial intelligence models in content generation, including at least contextual information related to the generation target such as information classification, target age, and story elements.

[0489] "Generative artificial intelligence models" refer to artificial intelligence models trained using machine learning or deep learning techniques that can automatically generate new text, images, or other content based on input text or other data.

[0490] "Prompt statements" refer to instructional statements that describe the generation conditions, content requirements, and style constraints in natural language or structured text. They are used as input to generative artificial intelligence models to guide them in generating expected output content.

[0491] "Story content" refers to text data generated by generative artificial intelligence models based on prompts, used to represent complete or partial storylines.

[0492] "Generated results" refers to all the data output by the generative artificial intelligence model after receiving prompts, including the story content text and related metadata.

[0493] "Imaging device" refers to an image acquisition device used to acquire images of a user's face or body, such as a built-in camera or an external camera.

[0494] "Voice capture device" refers to an audio input device used to capture the user's voice signal, such as a microphone or integrated audio capture module.

[0495] "Bioinformation" refers to signal data related to the user's state obtained by imaging devices and voice acquisition devices, including facial images, facial expression features, voice waveforms, and intonation features.

[0496] "Emotional analysis processing" refers to the process of analyzing biometric or textual information to infer the user's current emotional state, which typically includes steps such as feature extraction, pattern recognition, and emotion classification.

[0497] "Emotional state" refers to the state information obtained through emotion analysis that represents the user's psychological or emotional state, such as emotional labels or combinations thereof, such as happiness, sadness, surprise, tension, etc.

[0498] "Revision" refers to modifying, adding, deleting, or replacing parts of the story content according to specific rules, while keeping the overall structure of the story basically unchanged, in order to change its emotional tone, style, or details.

[0499] "Regeneration" refers to the process of using existing story content as one of the input conditions to regenerate new story content through a generative artificial intelligence model.

[0500] "Target emotional conditions" refer to the preset standards used to evaluate whether the emotional tendency of the story content is appropriate, including the target emotional type, the range of emotional intensity, or the degree of matching with the user's state.

[0501] "Rewriting prompts" refer to prompts used to instruct generative AI models to rewrite, refine, or regenerate existing story content, either entirely or partially.

[0502] "Book structure data" refers to structured data used to describe the internal page structure of an e-book or digital picture book, including information about multiple pages and the visual layout information associated with each page.

[0503] "Page information" refers to the structural unit data of a single page in the corresponding book, including at least the text content identifier of the page and related layout and visual information references.

[0504] "Visual information layout information" refers to layout data used to describe the spatial location, size, and hierarchical relationship of text areas, image areas, and other visual elements on each page.

[0505] "Image generation prompts" refer to text instructions used to guide artificial intelligence models in generating specific image content. These instructions typically include information such as scene descriptions, character characteristics, and style requirements.

[0506] "Artificial intelligence model for image generation" refers to an artificial intelligence model that can automatically generate corresponding image data based on input image generation prompts.

[0507] "Image data" refers to digital image information generated by an artificial intelligence model for image generation or obtained by other image acquisition devices, which is used for display on a page.

[0508] "E-book format" refers to a book presentation format that is stored as digital data and can be displayed on a terminal device by turning pages or scrolling, including electronic picture books, digital storybooks, and other formats.

[0509] In various embodiments of this invention, the server, terminal, and user each assume different functional roles. The server primarily performs computationally intensive processing such as data storage, feature extraction, prompt generation, generative artificial intelligence model invocation, sentiment analysis, regeneration control, and book structure data generation; the terminal primarily performs interface display, biometric information collection, and result presentation; and the user selects conditions and reads through the terminal. These embodiments of the invention can be implemented on general-purpose computer devices, such as server devices equipped with multi-core central processing units, graphics processing units, random access memory, and non-volatile memory, and connected to various terminal devices via wired or wireless networks.

[0510] I. Specific configuration of system hardware and software In one implementation, the server uses an operating system (e.g., a general-purpose server operating system) and runs a backend framework (e.g., a Python-based web framework) as an application server. The server further includes: - Storage device: Used to store story element candidates, user preference parameters, sentiment analysis model parameters, generative artificial intelligence model call logs, and book structure data. This storage device can be a relational database, document database, or cloud database service.

[0511] - Bioinformatics processing module: Used to process image and audio data sent from the terminal.

[0512] - Generative AI Model Interface Module: Used to communicate with text generation models (such as language models based on multi-layer Transformer architecture).

[0513] - Image generation model interface module: Used to communicate with image generation models (such as diffusion-based image generation networks).

[0514] - Sentiment Analysis Module: Used to call sentiment recognition services or local sentiment neural networks to perform sentiment analysis on biological information and story text.

[0515] - Prompt Statement Construction Module: Used to combine user conditions, story elements, and emotional states into structured natural language prompt statements.

[0516] - Book Structure Generation Module: Used to convert the generated story content into multi-page book structure data.

[0517] In one embodiment, the terminal is a smart mobile terminal or a tablet terminal, and the terminal includes: - Display device: Used to display the user interface and e-book pages.

[0518] - Touch input device: used to obtain the user's selection of information categories, target age and story elements.

[0519] - Camera device: Used to capture images of the user's face.

[0520] - Microphone: Used to capture the user's voice or sound during reading.

[0521] - Communication module: Used for bidirectional data communication with the server over a network.

[0522] - Rendering module: Used to generate page-turning display effects locally based on book structure data.

[0523] Users interact with the server through a terminal. Without needing to understand the internal algorithms and model structure, users can select categories, ages, and emotional preferences through a graphical interface and read storybooks or picture books generated by the system.

[0524] II. Structure and Training of Generative Artificial Intelligence Models and Emotion Models In a preferred embodiment, the server uses a generative artificial intelligence model employing a multi-layered self-attention language model architecture, where each layer includes a multi-head self-attention sub-layer and a feedforward neural network sub-layer. The server pre-trains and fine-tunes this language model offline. - The server uses a large-scale text corpus during the pre-training phase, employing an autoregressive language modeling objective with cross-entropy loss as the error function. The server updates the model weights through gradient descent and variant optimization algorithms (such as adaptive learning rate optimization).

[0525] During the fine-tuning phase, the server uses specially collected children's story corpora and graded reading corpora to further train the model in order to improve the generation quality in terms of story genre, language style and different age groups.

[0526] - The server employs data augmentation techniques during training, such as perturbing the order of story segments, replacing synonyms, and expanding sentiment tags, to improve the model's robustness to diverse prompts.

[0527] In one implementation, the server uses a sentiment analysis model, which can be a combination of convolutional neural networks (CNNs) and recurrent neural networks (RNNs), to process image and speech features. After receiving image frames from the terminal, the server extracts the facial region and then uses a CNN to extract facial expression feature vectors. For speech data, the server first performs a short-time Fourier transform to obtain a spectrogram, and then extracts speech emotion features using a CNN or a long short-term memory network. The server inputs these features into a fully connected layer and outputs a multi-dimensional emotion probability vector, such as a probability distribution of [happiness, sadness, surprise, tension, calm], and then uses the emotion label corresponding to the maximum value as the dominant emotion state.

[0528] The server can further use a text sentiment analysis model to perform sentiment analysis on the generated story text. This text sentiment model is also based on a neural network structure, establishing a mapping from text to sentiment labels through word embedding layers, encoding layers, and classification layers, with cross-entropy loss as the optimization objective.

[0529] III. Data Structures and Data Flow Between Modules Internally, the server uses structured data structures to represent the information passed between modules. For example: - The condition information object includes the following fields: information category (string), target age (integer), list of story elements (characters, scenes, plots, etc., strings), target emotion (enumeration), and user identifier (string).

[0530] - The prompt statement object includes the following fields: prompt text (string), maximum generated length (integer), and randomness parameter (floating-point number).

[0531] - The story content object includes fields: complete text (string), list of paragraphs (string array), and emotion tags (enumeration).

[0532] - The book structure data object includes fields: an array of pages, each page containing text content, image references, and layout parameters (including coordinates and dimensions).

[0533] Server modules use the aforementioned objects to make function calls and pass memory, avoiding unstructured string concatenation, thus facilitating error checking and performance optimization at the program level.

[0534] IV. Construction and Technical Effects of Prompt Statements In one implementation, the server constructs prompt statements in a structured manner, rather than simply concatenating them. The server performs the following processing in the prompt statement construction module: - The server selects different vocabulary levels and sentence complexity ranges based on the target age so that the prompts can clearly specify requirements such as "suitable for children aged X", "approximately Y words", and "use simple sentences".

[0535] - The server precisely lists character names, scene descriptions, and plot objectives in the prompts based on information categories and story elements, for example: Example 1: "Please create an adventure story for a 5-year-old child. The main characters are a curious little bear and its friend, a little bird. The story should include a plot about searching for a mysterious treasure in the forest. The overall atmosphere should be happy and lighthearted, and the story should be about 500 words long." Example 2: "The target audience is a 7-year-old child. The child is currently a little afraid, so please write a bedtime story to help him overcome his fear of the dark. The main character is a little star that shines, the ending should be warm and safe, and the language should be simple and easy to read aloud." Example 3: "Characters: Pirates; Setting: Mysterious Treasure Island; Plot: Solving one mystery after another to find the treasure. Please write a suspenseful but not scary adventure story for a 10-year-old child, about 800 words, including a clear beginning, development, climax and ending." By using structured and parameterized prompts, the server provides generative AI models with clearer constraints within the input space. This improves the consistency and controllability of the generated results with user conditions, reduces irrelevant or off-topic text, and ultimately enhances the efficiency of computing resource utilization. With the same hardware resources, the server can generate satisfactory stories with fewer calls and shorter generation lengths, achieving the technical effects of increased processing speed and reduced communication load.

[0536] V. Regenerative Control Driven by Emotional Feedback In one implementation, after obtaining the initially generated story content, the server inputs the story text into a text sentiment analysis model to obtain sentiment labels for the entire story or segments thereof. The server compares these sentiment labels with target sentiment conditions. If the preset conditions are not met, for example, if the target is "happy, warm," but the analysis result leans towards "neutral, slightly tense," the server performs the following processing: - The server constructs rewritten prompts, embedding the original story text within them, and explicitly instructs the generative AI model to adjust the story's emotional tone. For example: "Below is a story generated for an 8-year-old child, and the overall atmosphere is currently a bit tense. Please rewrite the story into a lighter, funnier version without changing the main characters and plot development, using simple Chinese and dialogue suitable for an 8-year-old child. The original story is as follows:..." The server inputs the rewritten prompts along with the original story text into a generative AI model. The model then re-encodes the input text using its internal attention mechanism to generate a new version of the story.

[0537] - The server performs sentiment analysis on the new story again. If the sentiment tags meet the target conditions, the version is used as the final story. If there are still deviations, the server can further tighten the constraints in the prompts, such as adding rules like "the ending must end with a celebratory scene," and then generate the story again.

[0538] Through this emotion-feedback-driven regeneration control, the server constructs a closed-loop optimization process within the computer, unlike the method of human text modification. This process relies on numerical modeling of the mapping between text features and emotion tags, enabling automatic and batch filtering and adjustment of large amounts of generated content, significantly reducing human intervention and improving the overall system's generation accuracy and consistency.

[0539] VI. Collaboration between Book Structure Data Generation and Image Generation In one implementation, the server divides the finalized story content into multiple pages, each consisting of one or several sentences. The server divides the content based on paragraph length and semantic boundaries to avoid truncating sentences in the middle of the page. In the book structure generation module, the server assigns layout parameters to each page; for example, text areas are located in the lower half of the page, and image areas are located in the upper half, recording coordinates and dimensions numerically.

[0540] The server automatically generates images and prompts based on the text content of each page. For example, when the page text describes "The little fox is playing with his friends in the forest," the server generates the following prompt: "Please generate a colorful illustration in the style of a children's picture book, showing a little fox playing with several animal friends in a green forest. The style should be cute and the colors bright." The server inputs the prompt into an AI model for image generation. This model employs a multi-layer convolutional network or diffusion model structure, iteratively optimizing the image from random noise to gradually approximate the potential target distribution encoded by the prompt in the feature space. This iterative process uses a specific loss function (e.g., a combination of perceptual and adversarial loss) to adjust model parameters or sampling paths, thereby generating illustrations that highly match the text description.

[0541] The server associates the image data obtained for each page with the corresponding page information and stores the text and images uniformly in the book structure data. During rendering, the terminal does not need to request text and image generation again; it only needs to draw the ebook page locally based on this structure data, thereby reducing the burden of subsequent communication and ensuring display consistency across different terminals.

[0542] VII. Technical Effects and Improvements in Computer Technology Through the aforementioned specific data structures, prompt statement construction rules, emotional feedback closed-loop control, and the collaborative process of book structure and multimodal generation, the server has achieved improvements in computer technology in the following aspects: - By using structured conditional information and prompts on the input side, the server makes the input distribution of the generative artificial intelligence model more concentrated, reduces the invalid search space, improves generation efficiency, and thus shortens the response time under the same hardware conditions.

[0543] - The server, through its sentiment analysis and regeneration control modules, performs automatic sentiment labeling and regeneration decisions after generating the output. This transforms the generation process from a one-time call into a feedback-enabled and adjustable control system within the computer. This numerical control logic, unlike the traditional simple model invocation approach, reduces the probability of the generated results deviating from the target, thereby improving the overall system output quality.

[0544] The server pre-generates page layouts and image references at the book's structural data level and transmits them to the terminal in a single transmission. The terminal then performs page-turning rendering and display locally. This method of completing typesetting and resource organization on the server side makes the terminal rendering process lightweight, reduces multiple small-granular requests, and lowers network communication load.

[0545] - Servers and terminals use a unified data model and module division, which allows the same data structures and algorithm processes to be reused in different hardware environments, making it easier to deploy, expand and maintain.

[0546] VIII. Other Implementation Forms and Alternative Solutions In other implementations, servers can employ different generative artificial intelligence model structures, for example: - Use an encoder-decoder sequence-to-sequence model to encode structured story element vectors and then decode them into text.

[0547] - By using a multimodal model, the sentiment feature vector is merged with the text input into the same attention network, allowing the model to consider both user sentiment and text conditions in a single forward pass.

[0548] In another implementation, the terminal can provide only text input without collecting biometric information. In this case, the server can omit sentiment analysis and generate story content solely based on information classification and the target age. Alternatively, the terminal can collect only images without collecting speech, and the server can infer the emotional state based solely on image features.

[0549] In terms of book structure generation, the server can automatically adjust layout parameters according to the screen ratio of different terminals. For example, for portrait-oriented terminals, the server sets a larger text area and a single-column display mode; for landscape-oriented terminals, the server sets a left-right split-column text and image layout mode. This adaptive layout processing is calculated based on the screen resolution and pixel density parameters reported by the terminal, thereby improving the reading experience on different devices.

[0550] In summary, by organically combining multiple modules such as generative artificial intelligence models, prompt statement construction, sentiment analysis, and book structure generation, the server implements a closed-loop, multimodal, and tunable story generation and presentation technology solution within the computer system using specific data structures and control logic. This not only partially replaces the human creative process but also improves the computer's processing efficiency and output quality through algorithm and architecture design.

[0551] use Figure 14 The processing flow is explained.

[0552] Step 1: The user selects conditions and initiates a request on the terminal.

[0553] Input: None (user's initial state).

[0554] Output: Request data containing information categories, target age, and optional story elements.

[0555] After launching an application or webpage on the terminal, users use touch controls to select information categories (e.g., "adventure" or "fantasy"), target ages (e.g., "5 years old" or "8 years old"), and can further select characters (e.g., "pirate" or "bear"), scenes (e.g., "forest" or "universe"), and plot elements (e.g., "treasure hunt" or "puzzle"). The terminal reads these selection values ​​from the interface components, performs type checks (e.g., converting the age field to an integer), and then encapsulates them into structured data (e.g., internal objects with key-value pairs) ready to be sent to the server.

[0556] Step 2: The terminal sends user condition data to the server.

[0557] Input: The condition object selected by the user in step 1.

[0558] Output: The request message sent to the server.

[0559] The terminal serializes the condition object into a request body (e.g., a JSON string), appends metadata such as session identifiers and timestamps locally, then establishes a network connection to the server via the communication module, constructs an HTTP POST request, places the serialized data into the request body, and sends it to the specified interface address on the server. After sending, the terminal records the request number for matching subsequent server responses.

[0560] Step 3: The server receives and parses the user's condition data.

[0561] Input: A request message from the terminal.

[0562] Output: The parsed condition information object.

[0563] After receiving the HTTP request from the terminal in the network interface module, the server reads the string in the request body into memory and calls the parsing function to deserialize it into an internal conditional information object. The server performs validity checks on the fields (e.g., checking if the information category is in a predefined list, and if the age is within a reasonable range). If a missing field is found, it fills in the default value (e.g., the default category is "adventure"), and finally generates a conditional information object containing the information category, target age, story elements, and user identifier.

[0564] Step 4: The server retrieves candidate story elements from the storage device.

[0565] Input: The condition information object from step 3.

[0566] Output: A set of candidate story elements that match the conditions.

[0567] After receiving the conditional information, the server uses the information category and target age in the conditions as query parameters and initiates a query operation to the storage device through the database access module. The server performs data operations, including filtering by category, filtering by age level, and sorting by tag rating, and compiles the matching character settings, scene settings, and plot templates into a candidate set. The server saves this set as a list object containing multiple elements, which will be used as input for subsequent random selection.

[0568] Step 5: The server randomly selects candidate story elements.

[0569] Input: The candidate set of story elements obtained in step 4.

[0570] Output: The combination of story elements used in this generation.

[0571] The server reads the candidate set in the random processing module, calls a random algorithm (such as a selection function based on a pseudo-random number generator), and selects at least one element from each type of element (character, scene, plot). The server can adjust the random distribution according to the target age (e.g., increase the weight of milder endings for younger users) and determine the specific elements through weighted random calculation. The server finally generates a story element combination object, which contains the character, scene, and plot description strings required for this generation.

[0572] Step 6: The terminal collects the user's biometric information and sends it to the server.

[0573] Input: The user's current facial expression and voice (physical signal).

[0574] Output: A bio-information data packet containing image and audio data.

[0575] While the user is reading or waiting for the results, the terminal activates its camera and microphone to periodically capture facial image frames and short audio clips. The terminal compresses the image resolution and downsamples and encodes the audio, generating appropriately sized image files and audio clips. The terminal encapsulates this data along with the user's identifier into a biometric data packet, which is then sent to the server via the communication module for sentiment analysis.

[0576] Step 7: The server performs sentiment analysis to infer the user's emotional state.

[0577] Input: Bioinformatics data packets from the terminal.

[0578] Output: A sentiment label or probability vector representing the user's current emotional state.

[0579] The server preprocesses images (such as face detection, cropping, and normalization) and performs short-time Fourier transform on audio to generate a spectrogram in its bioinformatics processing module. This data is then input into a pre-trained emotion neural network model. Within the model, the server performs feature extraction and forward propagation calculations to obtain the output probability of each emotion category. The emotion with the highest probability is selected as the dominant emotion state. The server associates this emotion label with a conditional information object for use in subsequent prompts and regeneration control.

[0580] Step 8: The server constructs the text prompt statement used for the initial generation.

[0581] Inputs: Conditional information object, story element combination object, user emotional state.

[0582] Output: Initially generated prompt text.

[0583] In the prompt statement construction module, the server combines information classification, target age, story elements, and target emotional requirements into a natural language description. The server sets a word count range and syntactic complexity based on age, and adds constraints such as "the overall atmosphere of the story should be warm and happy" or "help the child overcome their fear of the dark" to the prompt statement based on emotional state. The server concatenates this information into a continuous text using string templates and parameter mapping, for example: "Please create an adventure story for a 5-year-old child. The protagonist is a curious little bear and its friend a little bird. The story should include a plot of searching for mysterious treasure in the forest, with an overall happy and lighthearted atmosphere. The story length should be approximately 500 words." The server outputs this text as the core field of the prompt statement object.

[0584] Step 9: The server will input the prompt into the generative artificial intelligence model and obtain the initial draft of the story.

[0585] Input: The prompt text from step 8.

[0586] Output: Initial draft text of the story generated by a generative artificial intelligence model.

[0587] In the generative model interface module, the server takes the prompt as input and calls the generative AI model service over the network. The server provides the model with generation configurations such as maximum generation length and temperature parameters. Internally, the generative AI model performs forward propagation calculations using a multi-layered neural network, predicting and generating story text word by word based on the prompt. After receiving the results from the model, the server extracts the main text fields and saves them as the initial draft of the story. The server can perform simple data processing on the draft, such as removing redundant prefixes, cleaning up formatting, and splitting it into paragraphs.

[0588] Step 10: The server performs text sentiment analysis on the initial draft of the story.

[0589] Input: The initial draft text of the story obtained in step 9.

[0590] Output: Overall emotional label and emotional score of the story.

[0591] In the sentiment analysis module, the server inputs the story text, either in segments or as a whole, into the text sentiment analysis model, performs vectorized encoding and classification operations, and generates a probability distribution representing the story's emotional tendency. Based on the probability values, the server determines the main emotions of the story (e.g., "happy," "tense," "slightly sad"), and calculates the difference index (e.g., difference or distance measure) between the emotion and the target emotion conditions. The analysis result is then output for use by the regeneration control module.

[0592] Step 11: The server determines whether emotion-driven regeneration is needed and constructs rewritten prompt statements if necessary.

[0593] Input: The emotional tags and target emotional conditions of the first draft of the story, and the text of the first draft of the story itself.

[0594] Output: Rewrite the prompt text (or regenerate the unwanted decision results).

[0595] The server compares the emotional tags of the initial story draft with the user's target emotional conditions. If the difference is within an acceptable range, it outputs a "no need to regenerate" decision. If the difference exceeds a threshold, the server generates rewriting prompt text in the prompt construction module, for example: "Below is a story generated for an 8-year-old child. The overall atmosphere is currently slightly tense. Please rewrite the story into a lighter, funnier version without changing the main characters and plot development, using simple Chinese and dialogue suitable for an 8-year-old. The original story is as follows:..." The server embeds the original story text into the prompt, forming a new rewriting prompt.

[0596] Step 12: The server regenerates the story based on the rewritten prompts when needed.

[0597] Input: The rewrite prompt text from step 11.

[0598] Output: The final text of the story after emotional adjustment.

[0599] The server will re-input the rewriting prompts into the generative AI model, setting the corresponding generation parameters so that the model can rewrite the given story text. The generative AI model co-encodes the original story and the rewriting instructions, outputting a new story version. After obtaining this new text, the server can perform a brief sentiment analysis to confirm whether the target emotion has been achieved. If the conditions are met, the server marks this text as the final story content; otherwise, the server can further tighten the rewriting requirements or trigger regeneration again until the sentiment indicators meet the preset conditions or the generation limit is reached.

[0600] Step 13: The server converts the final story content into multi-page book structure data.

[0601] Input: The final text of the story after emotional adjustment.

[0602] Output: A book structure object containing information about multiple pages and layout data.

[0603] In the book structure generation module, the server groups the final story into several pages based on semantic segments or sentences, with each page corresponding to one or more sentences of text. The server calculates the text length and complexity of each page to ensure visual balance. For each page, the server generates a page information object, fills in text fields, and calculates layout parameters such as the position and size of text areas and illustration areas. The server then aggregates all page information into a book structure object, preparing for subsequent multimodal expansion.

[0604] Step 14: The server generates prompts for image generation on each page and calls the image generation model.

[0605] Input: The text content and layout requirements for each page.

[0606] Output: A set of image data matching each page and its reference information.

[0607] The server performs semantic extraction on the text content of each page, generates a summary description, and then combines this description with style requirements (such as "children's picture book style," "bright and cute colors") to form prompts for image generation. For example: "Please generate a colorful illustration in the style of a children's picture book, showing a little fox playing with several animal friends in a green forest. The style should be cute and the colors bright." The server sends these prompts to the AI ​​model for image generation through the image generation model interface module. The model internally performs image generation calculations and returns the corresponding image data. The server assigns a unique identifier to each image and associates it with the corresponding page information, recording it in the book structure object.

[0608] Step 15: The server sends the completed book structure data to the terminal.

[0609] Input: A book structure object containing text, image references, and layout information.

[0610] Output: The response message sent to the terminal.

[0611] The server serializes the book structure object, packages the page text, image resource locations, and layout parameters into response data, and sends it to the terminal as an HTTP response via the communication module. The server includes necessary caching control information in the response header so that the terminal can cache image resources as needed, reducing duplicate requests and thus lowering the load on subsequent communications.

[0612] Step 16: The terminal parses the book's structural data and generates a page-turning display interface.

[0613] Input: Book structure data sent by the server.

[0614] Output: An interactive e-book page displayed on the terminal.

[0615] After receiving the response, the terminal calls the parsing module to restore the structured data into a local page object array and loads the corresponding text and image resources. Based on the layout parameters, the terminal draws text and illustrations on the display device, generating the initial page content and setting up page-turning gesture listening logic. When the user swipes or clicks the page-turning button, the terminal switches the displayed content according to the page array index, re-rendering the next or previous page to achieve a continuous reading experience. The terminal can appropriately scale the layout parameters according to its own screen size to ensure reading comfort.

[0616] Step 17: Users can read stories on the device and provide feedback.

[0617] Input: The e-book page displayed on the terminal.

[0618] Output: Optional user feedback data (ratings, reviews, etc.).

[0619] Users read the text and illustrations on each page by observing the terminal screen. At the end or during reading, users can submit a satisfaction rating for the story using buttons on the interface, or enter brief comments such as "I wish the story were shorter" or "I wish there were more dialogue." The terminal collects this feedback data, encapsulates it into feedback objects, and sends it to the server in subsequent sessions. The server uses this data to update its internal preference parameters and prompt strategies, thereby further improving generation efficiency and matching accuracy in future projects.

[0620] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0621] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0622] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.

[0623] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0624] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.

[0625] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.

[0626] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.

[0627] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0628] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.

[0629] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0630] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0631] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0632] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0633] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0634] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0635] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.

[0636] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0637] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0638] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0639] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0640] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0641] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0642] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0643] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.

[0644] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0645] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.

[0646] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.

[0647] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.

[0648] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0649] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.

[0650] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0651] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0652] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0653] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0654] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0655] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0656] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.

[0657] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".

[0658] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0659] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0660] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0661] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0662] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0663] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0664] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.

[0665] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0666] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.

[0667] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.

[0668] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.

[0669] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0670] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.

[0671] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0672] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by a perspective equivalent to the field of vision of an average healthy person).

[0673] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0674] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0675] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0676] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0677] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0678] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.

[0679] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".

[0680] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0681] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0682] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0683] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0684] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0685] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0686] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.

[0687] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0688] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.

[0689] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see [reference]). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.

[0690] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.

[0691] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.

[0692] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).

[0693] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.

[0694] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."

[0695] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values ​​representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.

[0696] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).

[0697] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.

[0698] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0699] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.

[0700] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.

[0701] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.

[0702] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.

[0703] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.

[0704] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.

[0705] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.

[0706] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.

[0707] In addition, the following notes are provided in response to the above explanation.

[0708] Example 1 (Note 1) An information processing system, characterized in that it comprises: A processing unit running on an information processing device, the processing unit being configured to generate an interface screen for user operation via a display device connected to the information processing device, enabling the user to input content category information and target age group information, and to acquire selection information of the content category information and the target age group information; The system receives communication data containing the selection information from an information terminal that communicates with the information processing device, and dynamically generates prompt statements to instruct the generative artificial intelligence model on the conditions for generating content according to a predefined statement structure based on the type and value of the selection information. The system then sends the generation request data containing the prompt statements to the sending unit of the generative artificial intelligence model. The generative artificial intelligence model obtains the language data generated in response to the prompt statement, and the emotion recognition processing unit analyzes the expression content contained in the language data. The analysis result is correlated with the user's emotional state, and at least a part of the constituent elements contained in the language data is changed or added based on the emotional state to generate an adjustment unit of adjusted language data. The adjusted narrative data is converted into output data in an output format that can be displayed on the information terminal, and the output data is sent to the output unit of the information terminal via a communication link.

[0709] (Note 2) The information processing system according to Appendix 1 is characterized in that, The information processing device is configured to determine, based on the user's emotional state parsed by the emotion recognition processing unit, reconstruction conditions for reconstructing at least a portion of the main characters, scene structure, event progression, and ending structure constituted by the narrative data obtained from the generative artificial intelligence model, and to automatically re-edit the narrative data according to the reconstruction conditions, thereby generating narrative data that dynamically changes over time as the user's emotional state changes.

[0710] (Note 3) The information processing system according to Appendix 1 is characterized in that, The information processing device is configured to select multiple attribute information representing the story theme category, the main character category, and the expression style category from the storage device based on the content category information, the target age group information, and the user's emotional state identified by the emotion recognition processing unit, and attach the attribute information to the prompt statement as input conditions to the generative artificial intelligence model, so that the story data generated by the generative artificial intelligence model reflects the attribute information.

[0711] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: The information input / output device is used to obtain information categories and user attributes specified by the user from the terminal. Based on the acquired information category and user attributes, the system automatically generates prompt statements to instruct the generative artificial intelligence model to perform content generation, and sends the prompt statements in character sequence form to the prompt statement generation and sending unit of the generative artificial intelligence model. The generative artificial intelligence model receives generated content output sequentially in response to the prompt statement. During the receiving process, the generated content is acquired as segmented information and stored sequentially to form content data. At the same time, the segmented information is sent sequentially to the streaming unit of the terminal via the network in a streaming mode. In the terminal, a display control unit parses the segmented information received in a streaming manner and displays the segmented information sequentially on a visual display device through display control, so as to gradually present the generated content to the user; A content adaptation unit that analyzes the generated content to determine the expression difficulty and / or expression length corresponding to the user attributes, and adjusts the prompt statement and / or the generated content according to the determination result.

[0712] (Note 2) The information processing system according to Appendix 1 is characterized in that, The content adaptation unit extracts content elements by parsing the generated content into strings and / or symbols, and replaces and / or deletes some content elements based on the user attributes, thereby post-processing the generated content before sending it to the terminal.

[0713] (Note 3) The information processing system according to Appendix 1 is characterized in that, The prompt statement generation and sending unit adds at least one of the following to the prompt statement based on the information category and user attributes obtained through the information input / output device: word count condition, expression style condition, and / or plot progression condition. This controls the structure of the generated content output by the generative artificial intelligence model under the constraints of the added conditions.

[0714] Example 2 (Note 1) An information processing system, characterized in that it comprises: A display device and an input device for enabling users to input generation conditions, including the type of work and the target age, through an information processing device; The control device is used to parse the generation conditions obtained through the input device and automatically generate prompt statements for the generative artificial intelligence model based on the generation conditions, thereby instructing the generative artificial intelligence model to generate story data. An editing device for processing story data obtained from the generative artificial intelligence model, so that the story data is edited into page-divided work structure data based on vocabulary difficulty, sentence length, expression content and page division corresponding to the target age; A generation device for generating corresponding prompts for visual presentation based on the text content of each page in the work structure data, and associating the prompts as additional information for image generation or presentation with the work structure data. A recording device for storing the generation conditions and generation history of the work's structural data in a reusable form based on a user's save or regenerate request. An output device for sending story data processed by the editing device and prompts associated with the generating device to an information processing device in a page-turning format and displaying them thereon; An adjustment device for inputting generated story data into an analysis device with emotion recognition capabilities, and for adjusting the story's constituent elements according to the user's emotional state.

[0715] (Note 2) The information processing system according to Appendix 1 is characterized in that, The adjustment device is configured to automatically reconstruct at least a portion of the scene composition, characters, intensity of expression, and type of ending contained in the work structure data based on the user's emotional state identified by the analysis device, and to update the output of the generative artificial intelligence model sequentially, so that different users can obtain different story developments.

[0716] (Note 3) The information processing system according to Appendix 1 is characterized in that, The control device is configured to automatically select the theme of the work, the attributes of the characters, and the visual presentation guidelines based on the generation conditions and the user's emotional state identified by the parsing device, and reflect the selection results in the prompt statements for the generative artificial intelligence model, thereby coordinating the control of the story data generated by the generative artificial intelligence model and the visual presentation prompt statements for each page generated by the generation device.

[0717] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: A device for acquiring user operation information; A device for generating a user interface that allows a user to specify information categories and target age, and for outputting the user interface to a terminal device; An apparatus for retrieving story element candidates from a storage device based on the information classification and the target age, and for selecting at least one story element from the story element candidates through random processing; A device for automatically generating prompts for input into a generative artificial intelligence model based on conditional information including the information classification, the target age, and the story elements; An apparatus for inputting the prompt statement into the generative artificial intelligence model, causing the generative artificial intelligence model to generate story content, and obtaining the generation result; A device for performing emotion analysis processing based on biometric information acquired by an imaging device and a voice acquisition device to infer the emotional state of a user, and to modify or regenerate at least a portion of the story content based on the emotional state. A device for generating book structure data, including multiple page information and visual information layout information corresponding to each page, for the story content that has been modified or regenerated; A device for sending the book structure data to the terminal device and for displaying it on the terminal device as an e-book with page-turning functionality.

[0718] (Note 2) The information processing system according to Appendix 1 is characterized in that, The system is configured to: evaluate the emotional tendency of the story content in the emotional state inferred by the emotional analysis processing; when the emotional tendency does not meet the predetermined target emotional conditions, generate a rewriting prompt statement for re-input to the generative artificial intelligence model, wherein the rewriting prompt statement includes the story content, and control the generative artificial intelligence model to regenerate the story content based on the rewriting prompt statement.

[0719] (Note 3) The information processing system according to Appendix 1 is characterized in that, The system is configured to: when generating the book structure data, generate prompt statements for image generation based on the description content of each page after the story content is divided into pages, input the prompt statements for image generation into an artificial intelligence model for image generation to obtain image data, and associate the image data as visual information with the corresponding page information for layout.

Claims

1. An information processing system, characterized in that, Includes a processor, the processor being configured to: Provide users with a user interface for specifying genre and age; Based on the user-specified genre and age, automatically generate prompt text to instruct the generative AI model to perform story generation; The generated story is analyzed using an emotion engine designed to identify user emotions, and the story is then adjusted based on the analyzed emotions.

2. The information processing system according to claim 1, characterized in that, The processor is configured to cause the emotion engine to reconstruct the story elements of the generated story so that the generated story can dynamically change according to the user's emotions.

3. The information processing system according to claim 1, characterized in that, The processor is configured to automatically select themes and characters of a picture book based on the genre and age specified by the user and the identified emotions, and to reflect the themes and characters in a story generated by a generative artificial intelligence model.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A