Voice story playing method, device, equipment, storage medium and program product
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-14
- Publication Date
- 2026-08-11
AI Technical Summary
[0009]应当理解,本部分所描述的内容并非旨在标识本公开的实施例的关键或重要特征,也不用于限制本公开的范围。本公开的其它特征将通过以下的说明书而变得容易理解。
Smart Images

Figure CN122551767A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, specifically to intelligent speech generation, natural language processing and large language models, and in particular to a method, apparatus, electronic device, computer-readable storage medium and computer program product for playing voice stories. Background Technology
[0002] In recent years, the rapid development of speech synthesis technology has brought new opportunities to scenarios such as child companionship and storytelling. Current speech synthesis solutions mainly include applications based on traditional text-to-speech (TTS) technology. These methods generate high-quality speech through models, suitable for various occasions such as navigation announcements and audiobooks. Simultaneously, the introduction of speech cloning technology has made it possible to generate personalized voices, allowing users to create unique voice styles using collected speech samples. These technological advancements provide richer means for listening to stories and engaging in natural communication with children, making companionship interactions more vivid and interesting. Summary of the Invention
[0003] This disclosure presents a method, apparatus, electronic device, computer-readable storage medium, and computer program product for playing voice stories, which realizes real-time voice story generation with low latency, interruptibility, and rewriteability.
[0004] In a first aspect, embodiments of this disclosure propose a method for playing audio stories, comprising: extracting narrator identity parameters and listener identity parameters from a storytelling instruction input by a user; constructing identity setting values based on the narrator identity parameters and listener identity parameters; initializing a story state object based on the identity setting values, and generating streaming text segments based on the identity setting values and the real-time story state object; wherein the identity setting values are used as implicit information in the generation of each text segment each time; generating audio story segments from each sequentially generated streaming text segment according to the timbre information in the identity setting values and playing them to a target listener; wherein the target listener is determined based on the listener identity parameters.
[0005] Secondly, embodiments of this disclosure propose a voice story playback device, comprising: a parameter extraction unit configured to extract narrator identity parameters and listener identity parameters from a user-input storytelling instruction; an identity construction unit configured to construct identity setting values based on the narrator identity parameters and listener identity parameters; an object initialization unit configured to initialize a story state object based on the identity setting values and generate streaming text segments based on the identity setting values and the real-time story state object; wherein the identity setting values are used as implicit information in the generation of each text segment each time; and a voice generation unit configured to generate voice story segments from each sequentially generated streaming text segment according to the timbre information in the identity setting values and play them to a target listener; wherein the target listener is determined based on the listener identity parameters.
[0006] Thirdly, embodiments of this disclosure provide an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to implement the voice story playback method as described in the first aspect.
[0007] Fourthly, embodiments of this disclosure provide a non-transitory computer-readable storage medium storing computer instructions that enable a computer to implement the voice story playback method as described in the first aspect when executed.
[0008] Fifthly, embodiments of this disclosure provide a computer program product including a computer program that, when executed by a processor, can implement the steps of the voice story playback method as described in the first aspect.
[0009] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0010] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is an exemplary system architecture to which this disclosure can be applied; Figure 2 A flowchart of a voice story playback method provided in this embodiment of the disclosure; Figure 3 A flowchart illustrating a method for generating streaming text fragments provided in this disclosure embodiment; Figure 4A flowchart illustrating a method for generating streaming text fragments in an interruption scenario, provided by an embodiment of this disclosure; Figure 5 A flowchart illustrating a method for generating streaming text fragments in a multi-narrator scenario, provided as an embodiment of this disclosure; Figure 6 A flowchart illustrating a method for monitoring audience status provided in this disclosure embodiment; Figure 7a and Figure 7b A flowchart illustrating a method for playing a voice story in an application scenario provided by an embodiment of this disclosure; Figure 8 A structural block diagram of a voice story playback device provided in an embodiment of this disclosure; Figure 9 This is a schematic diagram of the structure of an electronic device suitable for performing a voice story playback method, provided as an embodiment of this disclosure. Detailed Implementation
[0011] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding; these should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description. It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0012] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0013] Figure 1 An exemplary system architecture 100 is shown, to which embodiments of the audio story playback methods, apparatuses, electronic devices, and computer-readable storage media of this disclosure can be applied.
[0014] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0015] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various applications for enabling information communication between the terminal devices 101, 102, and 103 and server 105 can be installed. These applications include voice interaction applications, search applications, and instant messaging applications.
[0016] Terminal devices 101, 102, and 103 and server 105 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with displays, including but not limited to smartphones, tablets, laptops, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices, and can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here. When server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When server 105 is software, it can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here.
[0017] Server 105 can provide various services through its built-in applications. Taking a voice-interactive application that can generate voice story segments based on storytelling commands as an example, when running this electronic photo album application, server 105 can achieve the following effects: First, extract the storyteller identity parameters and listener identity parameters from the storytelling commands input by the user; then, construct identity setting values based on the storyteller identity parameters and listener identity parameters; initialize the story state object based on the identity setting values, and generate streaming text segments based on the identity setting values and the real-time story state object; wherein, the identity setting values are used as implicit information in the generation of each text segment each time; finally, generate voice story segments according to the timbre information in the identity setting values for each sequentially generated streaming text segment and play them to the target listener; wherein, the target listener is determined based on the listener identity parameters.
[0018] It should be noted that, in addition to being obtained from terminal devices 101, 102, and 103 via network 104, storytelling instructions can also be pre-stored locally on server 105 through various means. Therefore, when server 105 detects that this data is already stored locally (e.g., when starting to process previously stored voice story generation tasks), it can choose to retrieve this data directly from locally. In this case, the exemplary system architecture 100 may also exclude terminal devices 101, 102, and 103 and network 104.
[0019] Since generating voice story segments based on storytelling instructions requires significant computing resources and power, the voice story playback methods provided in subsequent embodiments of this disclosure are generally executed by a server 105 with strong computing power and abundant computing resources. Correspondingly, the voice story playback device is also generally located within the server 105. However, it should also be noted that when terminal devices 101, 102, and 103 also possess sufficient computing power and resources, they can also perform the aforementioned calculations performed by the server 105 through their installed voice interaction applications, thereby outputting the same results as the server 105. Especially when multiple terminal devices with different computing capabilities exist simultaneously, but the voice interaction application determines that its terminal device has strong computing power and abundant remaining computing resources, it can allow the terminal device to perform the aforementioned calculations, thereby appropriately reducing the computing pressure on the server 105. Accordingly, the voice story playback device can also be located within terminal devices 101, 102, and 103. In this case, the exemplary system architecture 100 may also exclude the server 105 and the network 104.
[0020] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0021] Please refer to Figure 2 , Figure 2 A flowchart of a voice story playback method provided in this disclosure embodiment, wherein process 200 includes the following steps: Step 201: Extract the narrator identity parameters and audience identity parameters from the user's storytelling instructions; This step is intended for the entity executing the voice story playback method (e.g., Figure 1The server 105 shown extracts narrator and listener identity parameters from the user-inputted storytelling instructions. The storytelling instructions can be complete sentences spoken by the user via voice activation, such as "Dad will tell the baby a bedtime story," or the same content entered through a text interface. The executing entity can use automatic speech recognition technology to convert the user-inputted storytelling instructions into storytelling text, and then use natural language understanding technology to perform semantic role labeling and named entity recognition on the storytelling text to extract narrator and listener identity parameters. The narrator identity parameter refers to the subject who will perform the storytelling action, such as "Dad," "Mom," "Grandpa," or a user-defined nickname, which usually acts as the subject or initiator of the action in the sentence. The listener identity parameter refers to the audience to whom the story is intended, such as "Baby," "Treasure," or a specific child's name, which often acts as the object or indirect object of the prepositions "to" or "for." This embodiment utilizes natural language processing technology to improve the recall and accuracy of identity parameter extraction.
[0022] In this embodiment, the executing entity can determine the core predicate "tell" through syntactic analysis and find its corresponding subject and indirect object. Then, through role mapping, the identified noun phrases are mapped to an internally predefined set of identity tags. For example, if family members have registered identity tags such as "Dad," "Mom," and "Baby" during initialization, the extraction results will directly correspond to these tags. If the user uses an unregistered title (such as "Uncle"), it can be temporarily used as a custom identity tag, and the attributes of this identity can be improved in subsequent interactions through questioning or implicit learning. In practice, to improve the robustness of extraction, a few-shot parsing capability based on a large model can be introduced: by constructing a prompt template containing multiple instruction variations (e.g., "Please find the narrator and listener from the following sentences: {user instruction}"), the large model is called to output a structured JSON (JavaScript Object Notation) object, thereby flexibly handling complex sentence structures such as "Mom told me a story about Journey to the West" and "I want to hear Grandma tell her stories from her childhood," where the narrators are "Mom" and "Grandma," and the listeners are "I" (which needs to be further mapped to the current dialogue user) and empty (which can be set as the story-telling object by default).
[0023] In this embodiment, the extraction step does not require the user to explicitly and completely provide all identity parameters. When the narrator identity is missing in the instruction (e.g., simply saying "tell a story"), the executing entity can use the default narrator identity already locked in the current session (such as "Dad" determined in the previous conversation), or guide the user to complete it by actively asking. When the listener identity is missing, the current interactive user or the primary child user bound to the device can be defaulted to the listener. In addition, to cope with multi-user family scenarios, the extraction process can also be combined with voiceprint recognition or camera face recognition for auxiliary verification. That is, if the narrator identity parameter separated from the voice command matches the voiceprint features of the father already registered in the voiceprint database, the identity parameter is automatically supplemented or corrected, thereby improving the accuracy and naturalness of the extraction without increasing the user's verbal burden.
[0024] Step 202: Construct identity setting values based on the narrator's identity parameters and the audience's identity parameters; In this embodiment, the executing entity constructs identity setting values based on the narrator's identity parameters and the listener's identity parameters. Essentially, the identity setting value is a session-level state container that, in the form of key-value pairs or structured objects, solidifies the various attributes of the narration relationship into a set of computable and transitive variables. Specifically, the executing entity constructs the identity setting value by completing, mapping, and expanding the extracted parameters based on an internally predefined identity model. For example, when the extracted narrator's identity parameter is "Dad" and the listener's identity parameter is "Baby," the executing entity will associate "Dad" with its pre-registered timbre information, preferred self-reference (e.g., "Dad" instead of "I"), and typical tone of voice (e.g., gentle, slightly humorous) according to a family member configuration table or default rules, and associate "Baby" with corresponding terms of address (e.g., "Baby" or "Little Friend"). The generated identity setting value typically includes the following fields: a unique narrator identifier, narrator timbre information, narrator self-reference, listener address, language style vector or label, and optional emotional tone parameters.
[0025] In this embodiment, the core of constructing identity settings lies in parameterization and instantiation. Parameterization refers to converting role names in human language (such as "father") into enumeration values or IDs that can be indexed within the system. This process relies on a pre-established family member mapping table or dynamically registered voiceprint-identity binding relationships. Instantiation, on the other hand, involves retrieving the corresponding specific attribute values from the identity feature library based on these IDs and assembling them into a complete setting object. In practice, this construction operation is usually performed by an identity manager module, which maintains a session-level context object. Upon receiving identity parameters, the module first verifies whether the parameters are valid (e.g., whether the narrator has registered their voice). If necessary attributes are missing (e.g., the voice has not yet been cloned), the corresponding field can be set as a placeholder or default value, triggering an implicit data collection process in the background.
[0026] In this embodiment, to enhance the expressiveness and adaptability of the identity setting values, the construction process can also introduce rule-based or learning-based dynamic expansion mechanisms. For example, when the extracted audience identity parameter is "two children," the executing entity can automatically expand the audience address field to a plural form (such as "children") and increase the interactivity parameter in the tone style to adapt to the attention maintenance needs of the group audience. Furthermore, if the executing entity learns from historical interaction records that the audience prefers adventure themes or has a special interest in a certain animal character, it can proactively add an "interest tag" field to the identity setting values for use by the subsequent story material retrieval module.
[0027] Step 203: Initialize the story state object according to the identity setting value, and generate a streaming text fragment based on the identity setting value and the real-time story state object.
[0028] In this embodiment, the executing entity initializes the story state object based on identity settings. The story state object is a dynamically updatable, structured container designed for narrative generation, aiming to overcome the limitations of traditional dialogue systems that can only maintain short-term context or flat historical records. Initialization typically occurs at the beginning of the current session. The executing entity assigns initial fields to the story state object based on the narrator and listener characteristics implied in the identity settings, and the story theme clues carried in the storytelling instructions (if any). These fields include at least: a unique identifier for the current narrative node (usually starting from the root node), a summary vector representation of the generated text (initially empty or a preset opening summary), a character relationship status table (a basic character framework pre-defined according to the story type), plot progression markers (such as "beginning" or "developing"), and a session timestamp. After initialization, the story state object enters a ready state, awaiting entry into the generation loop.
[0029] In this embodiment, the executing entity generates streaming text fragments based on identity settings and real-time story state objects. The identity settings are implicitly included in the generation of each text fragment; that is, once constructed, the identity settings remain locked throughout the entire storytelling session unless explicitly initiated by the user. This means that even if the subsequent text generation process lasts for tens of minutes or spans multiple interaction rounds, the locked identity settings remain implicitly included in each generated text fragment, ensuring that the narrator's self-reference, address to the audience, and the synthesized voice tone remain consistent, preventing character drift. This implies that the input received during the generation of streaming text fragments consists of two parts: a dynamically changing story state object reflecting real-time information such as the current story progress, character relationships, and plot development; and a statically locked identity setting that fixes the narrator's character attributes, self-reference, audience address, tone of voice, and voice tone identifier. In practice, when a large language model is invoked for text generation, the executing entity converts each field in the identity setting into implicit information that the model can perceive. A common approach is to concatenate this implicit information into a system prompt. For example, the identity setting might be "You are a storyteller, you identify yourself as 'Dad,' you are telling a story to 'Baby,' and your tone should be gentle and humorous." This prompt is then input into the model along with the current story state object. Another approach is to use control vectors or adapter techniques to encode the identity setting into a set of embedding vectors. These vectors are then superimposed at each layer of the model during generation, thus influencing the selection of each output word with finer granularity.
[0030] In this embodiment, the process of generating streaming text fragments is cyclical: First, the executing entity reads the current story state object, extracts the narrative node identifier and summary vector to determine "what should be said next"; then, it merges the language style parameters in the identity setting with the narrative context, calls the text generation model to output a short, coherent text segment; after the text fragment is output, it is sent to the speech synthesis module for playback, and also sent back to the story state object to update the summary vector and advance the narrative nodes. This updated story state object then becomes the basis for the next round of generation. This process repeats until an interruption signal is detected or the user actively ends the session. Throughout the process, the identity setting always acts as a fixed "ballast," injected with equal weight in each round of generation, thus ensuring a high degree of consistency in the storytelling style.
[0031] Step 204: Generate audio story segments from each of the sequentially generated streaming text segments according to the timbre information in the identity settings and play them to the target audience.
[0032] In this example, the executing entity generates voice story segments from each sequentially generated streaming text fragment according to the timbre information in the identity settings and plays them to the target audience. Timbre information is a key field in the identity settings; it can be an identifier pointing to a pre-trained timbre model or a set of parameter vectors describing timbre features (such as fundamental frequency range, formant offset, spectral envelope parameters, etc.). In practical implementations, a common approach is to maintain a timbre library, where each registered narrator (e.g., "father," "mother") corresponds to a specific embedding vector in an independent neural speech synthesis model or a multi-speaker model. The timbre information field in the identity settings stores this embedding vector or model index, ensuring that all subsequent voice segments originate from the same vocal features. Traditional speech synthesis typically requires waiting for complete text input before synthesis begins, while streaming synthesis allows text fragments to arrive, be synthesized, and played simultaneously, significantly reducing end-to-end latency. Specifically, whenever the text generation module outputs a short text fragment (configurable in length, such as a complete sentence or dozens of characters), the fragment is immediately sent to the speech synthesis module's buffer. The synthesis module, based on the timbre information indicated by the identity settings, loads the corresponding acoustic model and vocoder, converting the fragment into Pulse Code Modulation (PCM) audio data. This audio data is then divided into smaller playback units (e.g., one frame every 10 milliseconds) and sequentially written to the audio output device. Simultaneously, the text generation module continues to produce subsequent fragments, forming a "generation-synthesis-playback" pipeline that ensures listeners experience virtually no pauses.
[0033] In this embodiment, the target audience is determined based on audience identity parameters. Playing to the target audience is not a physically directed broadcast operation, but a logical content adaptation. In actual use environments, smart speakers or voice terminals typically have only one main speaker, and everyone present can hear the generated sound. Therefore, playing to the target audience is more about the voice content itself being directed to a specific audience. For example, if the identity setting records the audience's title as "baby," the executing entity will frequently use this title in the generated and played voice (such as "Baby, guess what happened next?"), and the tone and word complexity will also be adjusted according to the audience's age and preferences. In multi-person family scenarios, the executing entity can also combine voiceprint locks or facial recognition to confirm whether the current main audience is the target audience specified by the identity setting. If not, it can prompt or wait for confirmation before continuing playback to avoid content mismatch.
[0034] The audio story playback method provided in this disclosure first extracts the narrator's and listener's identity parameters from user input commands and constructs a structured identity setting value. This identity setting value participates in the calculation as implicit information during the generation of each streaming text segment, avoiding the character drift problem common in long narratives. This ensures that the narrator's self-reference, tone, and emotional inclination remain highly consistent throughout the story playback, thereby providing users with a stable and reliable companionship experience. Then, the story state object initialized based on the identity setting value continuously records the current story progress, character relationships, and plot nodes, enabling the generation of subsequent text segments... Driven by the narrative state rather than relying on a flat historical record, this approach enhances the logical coherence of generating long or continuous stories, preventing plot breaks or inconsistencies. Finally, the collaborative approach of streaming text generation and streaming speech synthesis allows for immediate synthesis and playback of the first text fragment based on pre-associated timbre information set in the identity parameters, without waiting for the complete story content to be prepared. This significantly reduces the user's perceived initial response latency, ensuring a strong bond between the acoustic characteristics of the output voice and the narrator's identity. Even with long stories or multiple rounds of interaction, listeners receive a consistently personalized timbre experience. This embodiment, with identity locking and state-driven mechanisms at its core, enables voice interaction devices to possess low-latency, highly coherent, and identity-consistent real-time voice storytelling capabilities. It achieves real-time voice story playback with identity locking, state-driven mechanisms, and the ability to resume interrupted playback, significantly improving the user experience in intelligent companionship scenarios.
[0035] To further optimize the match between the story content and the identity portrayed, as well as the rationality of the narrative initiation, please refer to... Figure 3 , Figure 3 A flowchart of a method for generating streaming text fragments provided in this disclosure embodiment, wherein process 300 includes the following steps: Step 301: Determine the story type, initial scene, story characters, and narrative difficulty based on the identity setting values, and generate the initial story state object based on the story type, initial scene, story characters, and narrative difficulty; In this embodiment, the executing entity determines the story type, initial scene, characters, and narrative difficulty based on identity settings. The story type refers to the overall style or genre of the story, such as adventure, fairy tale, fable, science story, or heartwarming everyday story. The executing entity automatically matches the most suitable type based on the narrator's characteristics and the audience's age in the identity settings. For example, when the narrator is a "father" and the audience is a "toddler," the executing entity tends to choose a simple fairy tale or everyday story with positive character education; while when the audience is a "school-aged child," it may choose an adventure or science story. The initial scene is the first specific environment in which the story takes place, such as a forest, castle, family living room, or outer space. Its selection must be consistent with the story type and the audience's cognitive level implied in the identity settings. The characters are the set of characters that will appear in the story, typically including protagonists, supporting characters, and potential villains. The executing entity generates approachable or interesting character names and images based on the audience's age and interests. Narrative difficulty is a comprehensive parameter that controls the sentence length, vocabulary complexity, number of plot layers, and frequency of plot twists in the story. The higher the difficulty, the more complex the text, which is more suitable for older audiences.
[0036] In this embodiment, the executing entity can maintain a configuration table or decision tree, which records the default preferences corresponding to different narrator and audience combinations. For example, if the narrator in the identity setting refers to himself as "Dad" and the audience refers to him as "Baby," and the audience's age tag is 3 to 5 years old, then the rules can be set as follows: the story type is "Fairy Tale" or "Life Habit Story," the initial scene is "Little Bear's Home" or "Forest Kindergarten," the story characters are animals familiar to young children such as "Little Bear, Little Rabbit, Little Cat," and the narrative difficulty is the lowest level (no more than 10 words per sentence, no more than 3 sentences per paragraph). If the identity setting also carries user history preferences (such as previously requesting dinosaur stories), the story type can be dynamically adjusted to "Dinosaur Adventure." These determined parameters are passed to a story state object initializer, which assigns a unique session identifier, creates a blank role relationship state table (filling in the above story characters, with each character's initial state such as "healthy" or "not yet appeared"), sets the identifier of the first narrative node (e.g., "Node_001: Opening"), and marks the plot progression as "Beginning." Simultaneously, the descriptive text of the initial scene is encoded or directly stored in the corresponding field of the state object, serving as prompt material when generating the first subsequent text fragment. Narrative difficulty parameters are recorded to control output constraints during future streaming generation (e.g., limiting maximum generation length or controlling vocabulary levels). In this way, the executing entity can deduce a targeted and structurally complete initial story state solely based on identity settings, thus laying a solid and consistent foundation for subsequent streaming generation. This embodiment does not require users to provide additional story outlines or complex parameters; it relies entirely on intelligent decision-making driven by identity settings, significantly reducing the interaction burden and improving the relevance of content to the audience.
[0037] Step 302: Use the preset large model to determine the language style corresponding to the identity setting value and the story material corresponding to the current story state object; In this embodiment, the executing entity uses a pre-set large model to determine the language style corresponding to the identity setting value and the story material corresponding to the current story state object. The language style refers to the tone, wording, and sentence structure the narrator should use to organize their language, driven by the identity setting value. The story material refers to the specific plot content, character behavior, or scene description that should be narrated at the current story progress, driven by the current story state object. The pre-set large model refers to a generative language model pre-trained on a large-scale text corpus, such as a general model based on the Transformer architecture, which has the ability to understand complex instructions and output structured or unstructured text. When determining the language style, the executing entity does not directly have the large model output a style description, but rather uses an instruction-guided approach: serializing key fields in the identity setting value (such as the narrator's self-identification, audience address, tone style label, narrative difficulty parameters, etc.) into a natural language prompt. For example, construct a prompt template: "You are a storyteller, you identify yourself as 'Dad,' you are telling a story to 'Baby,' your tone should be gentle and fun, and the sentences should be short." After inputting this prompt into a large model, the first paragraph of text output by the model actually implicitly contains the style tendency for its subsequent generation.
[0038] In this embodiment, the story state object stores information such as narrative node identifiers, summary vectors of generated text, character relationship status tables, and plot progression markers. The execution entity transforms this information into a form understandable to the large model: mapping narrative node identifiers to a plot description (e.g., "Current node is 'Encounter in the Forest,' character A is on their way to the castle"), organizing the character relationship status table as "Character A is a kind little bear, character B is a cunning fox," and representing the plot progression marker as "The story is in the development stage and a small conflict needs to be introduced." This content, along with the language style in the identity settings, constitutes the input context of the large model. The story material generated by the large model can be directly usable text fragments or a structured plot blueprint (e.g., "Next, the little bear encounters a lost little bird in the forest"), for subsequent modules to refine. To balance effectiveness and efficiency, the implementing entity can pre-build a story material index library or knowledge graph. The large model first recalls the most relevant candidate materials (such as similar plot fragments or character dialogue templates) from the library through retrieval enhancement. Then, it combines the recalled materials to perform content fusion and rewriting, thereby generating specific content that is both in line with the narrative progress and has a sense of novelty.
[0039] Step 303: Generate streaming text fragments based on language style and story material.
[0040] In this embodiment, the executing agent generates streaming text fragments based on language style and story material. Specifically, the executing agent typically represents language style as a set of control parameters or a set of prompts, and story material as plot points to be developed or entity relationships to be described. These two parts are then concatenated into a complete model input, and the larger model generates streaming text fragments based on this input. For example, the input could be constructed as: "[Style Requirements] You are a narrator who calls himself 'Dad,' with a gentle tone and short sentences. [Material Content] A little bear encounters a lost bird in the forest. Please describe this encounter scene." After receiving this input, the text generation model begins outputting the first text fragment, such as, "The little bear was walking when it suddenly heard a soft crying sound coming from a tree branch. It looked up and saw that a little bird had gotten lost."
[0041] In this embodiment, streaming generation means that the model does not output complete paragraphs or chapters all at once. Instead, it uses autoregressive decoding to generate one word or sub-word unit at a time and immediately pushes the unit to the output buffer. To control segment boundaries, the executing entity can set hard conditions for segment termination, such as submitting the currently output content as a complete text segment to the downstream speech synthesis module when a period, question mark, exclamation mark, or newline character is encountered; alternatively, a fixed number of tokens can be set, with each N tokens generated considered as a segment. The former method produces segments with more complete semantics, facilitating natural pauses in synthesized speech; the latter method has a more uniform delay and is simpler to implement. Regardless of the segmentation strategy used, the essence of the generation process is to cyclically execute the sequence of "sample the next word - determine the boundary - output the segment - update the context - continue sampling" until a preset stopping condition is detected (e.g., completing a complete scene description or receiving an interruption signal).
[0042] In this embodiment, after each text fragment is output, its content is fed back to update the summary vector and narrative node in the story state object, rather than being simply discarded. This allows the model to adjust subsequent expressions based on the already "spoken" content when generating subsequent fragments, avoiding repetition or contradictions. Simultaneously, since language style constraints are repeatedly injected with each generation call (rather than just the first injection), the output intonation, self-identification, and narrative difficulty remain stable even if the generation process spans multiple fragments. To reduce the computational overhead of repeated injection, a caching mechanism can be employed, namely, encoding the language style part as a prefix key-value cache. This allows the model to reuse the hidden state of this part when processing different story materials, thereby improving generation throughput without sacrificing style consistency. Through the above methods, this embodiment achieves controlled, incremental text output synchronized with the narrative state in real time, providing a high-quality, low-latency input stream for subsequent speech synthesis and playback.
[0043] The method for generating streaming text fragments provided in this embodiment first proactively determines the story type, initial scene, characters, and narrative difficulty based on identity settings. This adaptive planning of the overall story framework before generation ensures the generated story aligns with the dual characteristics of the narrator and the audience from the outset, avoiding the stiffness of generic content. Then, a pre-defined large model is used to determine the language style corresponding to the identity settings and the story material corresponding to the current story state object, achieving decoupling between character tone and narrative content. Finally, streaming text fragments are jointly generated based on the language style and story material, ensuring that each output fragment carries the narrator's identity imprint in its word choice and faithfully reflects the current narrative progress and character relationships in its plot. This embodiment, through the above implementation, improves the fit between the story's opening and identity settings, avoids monotonous openings, enhances the flexibility and controllability of content generation, and matches the complexity of the output text with the audience's cognitive level, improving acceptance and immersion in companionship scenarios.
[0044] Based on any of the above embodiments, in order to achieve semantic-level interactive response and content continuation capability when an interruption event occurs, please refer to... Figure 4 , Figure 4 A flowchart of a method for generating streaming text fragments in an interruption scenario provided by an embodiment of this disclosure, wherein process 400 includes the following steps: Step 401: In response to detecting an interruption message from the target audience during playback, save the currently played streaming text segment as an intermediate text segment; In this embodiment, when an interruption message from the target audience is detected during the playback of the streaming text, the executing entity saves the currently played streaming text segment as an intermediate text segment. The interruption message refers to any signal that the system can recognize as a desire to pause, alter, or interfere with the current storytelling process, such as questions about the story content like, "Wait a minute," "I have a question," "Tell me again," or "Wait a minute, what was that character's name?" or "Why did he do that?" The executing entity needs to continuously run a parallel monitoring module, which continuously performs activity detection and keyword recognition on the audio captured by the microphone during voice playback, or polls other sensor signals, in order to capture the interruption message with the shortest possible delay.
[0045] In this embodiment, the saved intermediate text fragments refer to the text content that has been fully output through the speaker and actually heard by the listener before the interruption information is detected. It does not refer to text that has been generated but not yet synthesized into speech, nor to the text corresponding to audio that has been synthesized but is still queued in the playback buffer. To obtain accurate playback boundaries, the audio output module needs to return an acknowledgment callback to the control module upon completion of each audio block (e.g., every 20 milliseconds of PCM data). The callback carries the start and end positions of the original text corresponding to that audio block. The execution entity maintains a monotonically increasing playback pointer. When an interruption signal is received, it immediately reads the text position currently pointed to by the pointer and extracts all text from the beginning of the fragment to this position, storing it as an intermediate text fragment in a temporary storage area. This intermediate text fragment is essentially a "snapshot" of the narrative content; it records the precise endpoint of the story progression received by the listener. Subsequent state updates and restoration generation need to be based on this precise historical anchor point. For example, if the audience asks to "repeat the previous paragraph", the executing agent needs to know the text boundaries of "the previous paragraph"; if the audience asks a question about the events that just happened, the executing agent needs to refer to this text to understand the object referred to in the question.
[0046] Step 402: Use the large model to perform intent parsing on the interruption information and generate interruption semantic information; In this embodiment, the executing entity uses a large model to parse the intent of interruption information and generate interruption semantic information. Interruption information is typically short and contains omissions, colloquial expressions, or emotional connotations. The executing entity combines the ambiguous interruption information with a small amount of context from the current session to form a parsing prompt, which is then input into the large model. The large model outputs structured interruption semantic information according to a specified format. For example, for the interruption information "What's the lion's name?", the model can output a JSON object containing the intent type "asking about character attributes," the target character "lion," the required attribute "name," and the time anchor "in the current plot." For the interruption information "tell it again," the intent type is "repeated narration," and the repetition scope defaults to "the last paragraph." For the interruption information "I'm a little scared," the intent type is "emotional expression," with an attached emotion tag "fear." The executing entity can then decide whether to provide plot reassurance or switch the narrative. Therefore, the generated interruption semantic information is a structured data object containing intent category, focus, parameters, and confidence level, facilitating accurate execution by the subsequent state update module.
[0047] Step 403: Update the current story state object based on the intermediate text fragments and interruption semantic information, and generate a streaming text fragment based on the updated story state object and identity setting values.
[0048] In this embodiment, the executing entity updates the current story state object based on intermediate text fragments and interruption semantic information, and generates streaming text fragments based on the updated story state object and identity settings. Specifically, the executing entity combines the narrative boundaries already received by the audience (marked by intermediate text fragments) with the user's interruption intention (carried by interruption semantic information), modifies fields in the story state object that record narrative nodes, character relationships, plot progression, etc., so that the updated story state object can accurately reflect the state of the story world after the interruption event. After the story state object is updated, the originally locked identity settings are reused, and the new story state object is used as the content driver to restart the streaming text generation loop, thereby producing subsequent content that can respond to the user's interruption request.
[0049] The method for generating streaming text fragments in interruption scenarios disclosed in this embodiment first saves the currently played streaming text fragment as an intermediate text fragment when an interruption message from the target audience is detected during playback, thereby accurately recording the narrative boundaries already received by the audience; then, a large model is used to perform intent parsing on the interruption message and generate structured interruption semantic information; finally, the current story state object is updated based on the intermediate text fragment and the interruption semantic information, and the streaming text fragment is regenerated based on the updated story state object and the originally locked identity setting value, so that the subsequent output content can not only naturally respond to the interruption request, but also maintain the consistency between the narrator's identity and the story logic, thereby improving the degree of interactive freedom and narrative coherence during the audio story playback process, enabling smart devices to have the ability to respond naturally and flexibly continue the story, similar to a real person narrating.
[0050] Based on any of the above embodiments, to enrich the artistic expression and companionship of voice stories, please refer to... Figure 5 , Figure 5 A flowchart of a method for generating streaming text fragments in a multi-narrator scenario provided by embodiments of this disclosure is included, wherein process 500 includes the following steps: Step 501: In response to extracting at least two narrator identity parameters from the storytelling instruction, construct a multi-identity setting value based on the at least two narrator identity parameters and the audience identity parameters; In this embodiment, when at least two narrator identity parameters are extracted from the storytelling instruction, the executing entity constructs a multi-identity setting value based on the at least two narrator identity parameters and the audience identity parameter. The narrator identity parameters include a character identity identifier and the corresponding character voice information. Unlike the single-narrator scenario, the multi-identity setting value is no longer a single character object, but an ordered or unordered list of characters, where each item independently records the corresponding narrator's identity identifier and voice reference. The audience identity parameter is the same as in the single-narrator scenario, still pointing to the story's target audience.
[0051] In this embodiment, the executing entity first lists the extracted narrator identity parameters, for example, obtaining an array [{role: "Dad", voice ID: "voice_dad"}, {role: "Mom", voice ID: "voice_mom"}]. Then, the executing entity adds some metadata about collaborative narration to this list. For example, the "narration order" field can be set to "rotation", "free", or "character dialogue mode". Rotation mode means that the executing entity will assign the narration role to each text segment sequentially according to the list order. Free mode allows the large model to dynamically determine which narrator should speak in the current segment based on the needs of the plot. Character dialogue mode binds each narrator to a specific NPC character in the story, automatically calling the corresponding voice when the NPC speaks. Simultaneously, the multi-identity settings can also include collective self-reference parameters, such as "Dad and Mom" as a collective term for the overall narration, and a unified address for the audience.
[0052] In this embodiment, when the user only says "Let Mom and Dad tell the story," but does not specify who the first narrator is, the executing entity can default to setting the initial narrator according to the extraction order or the priority order of family members. If a narrator's voice modeling is not yet complete (for example, Dad's voice is still being implicitly acquired), the executing entity can temporarily use the default voice as a substitute and mark that character's voice as "to be improved" in the multi-identity settings, which does not affect the start of the multi-identity story. In addition, the multi-identity settings will also be locked during the session lifecycle unless the user explicitly changes or adds or removes the narrators. This means that once the setup is complete, all subsequent interruption recovery and story state object updates will retain this multi-role configuration, and the existence of other narrators will not be lost due to the generation of a certain segment. Through the structured multi-identity settings, a data foundation is laid for subsequent story type selection, character relationship table construction, and alternating multi-voice narration, enabling the voice story to present a richer performance level.
[0053] Step 502: Determine the story type, initial scene, story characters, and narrative difficulty based on the multiple identity settings, and generate the initial story state object based on the story type, initial scene, story characters, and narrative difficulty; In this embodiment, the executing entity determines four narrative dimensions—story type, initial scene, story characters, and narrative difficulty—based on the narrator character list and audience identity parameters included in the multi-identity setting values. Then, it generates an initial story state object based on the story type, initial scene, story characters, and narrative difficulty. Unlike single-narrator scenes, the multiple character identity identifiers in the multi-identity setting values directly influence the choice of story type and the composition of story characters. For example, when the narrator list contains "Dad" and "Mom," the executing entity may tend to choose a heartwarming family adventure or cooperative puzzle story type; when the narrator list contains "Sun Wukong" and "Tang Sanzang," the story type is more likely to lean towards mythology or a master-disciple adventure. The selection of the initial scene also needs to accommodate the appearance of multiple characters. The executing entity will select a location that can accommodate all the main narrator avatars appearing simultaneously or taking turns, such as "in front of the Water Curtain Cave of Flower Fruit Mountain" or "the living room at home." The determination of story characters is more direct; each narrator character in the multi-identity setting values can be directly mapped to a corresponding character in the story. Simultaneously, the system will generate several additional supporting characters to enrich the plot. The narrative difficulty needs to take into account the age of the audience and the complexity of the narrator's role. If there are multiple characters with very different styles (such as one serious and one humorous), the narrative difficulty may be increased accordingly in order to better show the contrast and interaction between the characters.
[0054] In this embodiment, the executing entity can pre-maintain a multi-role narrative preference table, which records the default story type and initial scene template corresponding to different role combinations. When the multi-identity setting is [{role: "father"}, {role: "mother"}], the executing entity finds that the "couple combination" preference is "family life" or "cooperative adventure", and the initial scene can be "forest cabin" or "beach". If the user instruction specifies an additional story theme (such as "tell a story about dinosaurs"), the story type will be preferentially adapted to the user theme. After determining the story type and initial scene, the executing entity writes each role identity identifier in the multi-identity setting as a required story character into the initial record of the role relationship status table, and sets an initial state for each role (such as "appeared at the beginning of the story"). The narrative difficulty can be set by combining the age tag in the audience identity parameter with the average language complexity of the narrator role, ensuring that the generated text is both understandable to the audience and reflects the personality characteristics of multiple characters.
[0055] When generating the initial story state object, the executing entity instantiates an empty state container and then fills in the fields determined by the four dimensions mentioned above. Specifically, the story type is encoded as a type ID and stored in the metadata area of the state object; the descriptive text of the initial scene (e.g., "On a sunny morning, Mom, Dad, and Little Bear were taking a walk in the forest") serves as the starting content cache for the first narrative node; the list of story characters is expanded into a character relationship state table, with each character occupying one row, recording their name, type (main narrator / supporting character), current status (e.g., "healthy" or "active"), and relationship with other characters; the narrative difficulty is stored as an integer value or enumeration, used to control sentence length and vocabulary complexity during subsequent streaming text generation. After initialization, the story state object possesses a complete starting context, providing a clear narrative starting point for subsequent material retrieval and text generation.
[0056] Step 503: Use the preset large model to determine the language style corresponding to the multiple identity settings and the story material corresponding to the current story state object; In this embodiment, the executing entity uses a pre-defined large model to determine the language style corresponding to multiple identity settings and the story material corresponding to the current story state object. Unlike the single narrator scenario, the language style corresponding to multiple identity settings is no longer a single, global tone description, but a style set that includes the expressive characteristics of multiple characters. For example, when the multiple identity settings include two characters, "father" and "mother," the large model needs to infer a "calm and slightly humorous" style for father and a "gentle and slightly fast-paced" style for mother; when "Sun Wukong" and "Tang Sanzang" are included, the corresponding styles might be "lively and mischievous" and "solemn and patient." The executing entity can serialize the multiple identity settings into a character list description and construct the following prompt: "Your task is to determine the appropriate language style for each character based on the following list of narrator characters. Character 1: Father, Character 2: Mother. Please output the style keywords for each character in JSON format." The large model generates a structured style mapping accordingly, which is dynamically switched based on the currently active character during subsequent text generation.
[0057] In this embodiment, the executing entity uses the same or another large model call to determine the story material corresponding to the current story state object. The mechanism for determining the story material is similar to that in a single-narrator scenario, that is, based on the narrative nodes, character relationship table, plot progress markers, and generated summary vectors stored in the story state object, it retrieves or generates the plot content to be told next. However, in a multi-narrator scenario, the determination of the material also needs to consider which narrator is currently speaking. The story state object usually records a "currently active narrator" field, and the executing entity incorporates the timing of character switching into the material generation, so that the material content naturally matches the character's personality. For example, when it is the father's turn to narrate, the generated material may contain more action-oriented descriptions; when it is the mother's turn, more emotional expressions or parent-child interactions may be inserted. The large model can simultaneously receive multi-identity style mappings and story state objects, and output a material description that conforms to the narrative progress and reflects the current character's language characteristics. This material can be a summary of plot points (such as "Next, the little bear got lost, and the father told the little bear not to be afraid in an encouraging tone"), which is then handed over to the subsequent refinement generation module to expand into specific text fragments.
[0058] In this embodiment, to improve efficiency, language style determination can be completed and cached all at once during session initialization, as the multiple identity settings remain unchanged throughout the session. Story material determination, however, is re-executed with each update of the story state object (including interruption recovery). Regarding the choice of the main model, the same model instance can be used for both sub-tasks, but different prompt templates can be used to distinguish the task types. If latency needs to be reduced, the language style determination task can be handled by a lightweight rule matcher or a low-parameter model, while the story material determination task can be reserved for the main model.
[0059] Step 504: Generate streaming text fragments based on language style and story material.
[0060] Step 504 and as follows Figure 3 The steps shown in step 303 are the same. For the same parts, please refer to the corresponding parts of the above embodiments, which will not be repeated here.
[0061] The method for generating streaming text fragments in a multi-narrator scenario disclosed in this embodiment extracts at least two narrator identity parameters from the storytelling instruction and constructs a multi-identity setting value containing multiple character identity identifiers and their corresponding timbre information. This setting value enables the subsequent initialization of the story state object to simultaneously consider the existence of multiple narrators. The story type, initial scene, story characters, and narrative difficulty determined accordingly are naturally adapted to the collaborative mode of multiple narrators. When using a large model to determine the language style, differentiated tone expressions can be generated for different characters. Furthermore, the text fragments output by each character are synthesized with their bound exclusive timbre. When the story content involves dialogue or character switching, different timbre models can be alternately called for broadcasting based on the order in the multi-identity setting value or the decision of the large model, thereby achieving an auditory effect similar to role-playing reading or multi-person co-narration, enhancing the vividness and immersion of storytelling, and enabling listeners to clearly identify different characters through voice differentiation, providing richer expressive means for parent-child interaction or group companionship scenarios.
[0062] Based on any of the above embodiments, to further enhance context awareness and user engagement, please refer to... Figure 6 , Figure 6 A flowchart of a method for monitoring audience status provided in this disclosure embodiment, wherein process 600 includes the following steps: Step 601: Monitor the audience's state using image acquisition equipment and / or voice acquisition equipment; In this embodiment, to detect changes in the audience's state and subsequently decide whether to initiate active interaction or adjust content, the executing entity uses image acquisition devices and / or voice acquisition devices to monitor the audience's state. Image acquisition devices typically refer to cameras or infrared sensors integrated into voice terminals (such as smart speakers or companion robots) to capture the audience's visual information; voice acquisition devices refer to microphone arrays to capture the audio signals emitted by the audience. These two devices can be used independently or in combination to improve the accuracy and robustness of the monitoring. The monitored audience state includes at least attentional state (such as whether they are looking in the direction of the device, yawning, or turning their heads) and emotional state (such as whether they exhibit impatient humming, laughter, or crying sounds). Changes in these states are collectively referred to as candidate signals of attentional shift or abnormal states.
[0063] In this embodiment, to achieve low-power real-time monitoring, the execution entity typically does not continuously run highly complex visual models, but instead employs a hierarchical detection mechanism. The first level uses motion detection or speech energy detection to determine if the listener is near the device; if someone is detected, a lightweight face detection and keypoint analysis model is periodically activated (e.g., 2-5 frames per second). For speech monitoring, a low-complexity wake-word-independent speech activity detector can be continuously run, triggering the emotion recognition model only when a non-story speech segment is detected. All monitoring results are recorded in the session context with timestamps for subsequent steps to query. Furthermore, to protect user privacy, the image acquisition device can process and discard the original image locally in real time, outputting only abstract state features (e.g., "attention shift confidence 0.8"); the speech acquisition device can also perform acoustic feature extraction and classification locally without uploading the original audio. Through the above technical means, this embodiment can continuously acquire listener state information in a low-latency, low-power manner while protecting privacy, providing reliable input for subsequent proactive interaction strategies.
[0064] Step 602: In response to the detection of a shift in audience attention, a streaming text segment is generated based on a preset interactive template; In this embodiment, when an attention shift is detected in the audience, the executing entity generates a streaming text segment based on a preset interaction template. An attention shift refers to a quantifiable deviation from the audience's initial focus on the story content. For example, this could be detected by an image capture device as the audience's gaze moves away from the device, their head continuously turns, or their eyes remain closed for an extended period; or by an audio capture device as the audience humming, sighing, or conversing with a third party without regard to the surrounding environment. Once these state characteristics are confirmed as attention shifts by preset judgment rules (such as a deviation duration exceeding two seconds or a confidence level exceeding a threshold), the response logic of this step is triggered.
[0065] In this embodiment, monitoring based on image acquisition devices typically employs target detection and behavior recognition technologies from computer vision. The executing entity first locates the listener's facial region using a face detection algorithm, and then extracts features such as eye gaze direction, eye opening and closing degree, and mouth shape using key point detection (such as key points of the eyes and mouth). By combining changes over multiple consecutive frames, it can determine whether the listener is looking at the speaker, whether their eyes are closed (possibly due to drowsiness), whether they are yawning, whether they are turning their head away, or whether they exhibit facial expressions of fear / frustration. For example, when the listener's head deflection angle exceeds 30 degrees and the duration exceeds a threshold, or when the eye closure time significantly increases, it can be determined as an attention shift. Monitoring based on voice acquisition devices relies on voice activity detection and emotion recognition technologies. The executing entity continuously acquires ambient audio, excludes the story voice played by the speaker itself through voiceprint separation, analyzes the non-verbal sounds (such as sighs, humming) or short spoken words (such as "bored" or "I'm not listening") emitted by the listener, extracts acoustic features such as fundamental frequency, energy, and speech rate, inputs them into a pre-trained emotion classification model, and outputs attention or emotion state labels. These two monitoring methods can complement each other. Image devices are highly accurate when the light is sufficient and the listener is facing the device, but they fail when the listener is facing away or in low light. Voice devices are not affected by light obstruction, but may be affected by environmental noise. Therefore, in actual products, a fusion strategy is often adopted, that is, when the confidence of one signal is low, the other signal is referenced for comprehensive judgment.
[0066] In this embodiment, the preset interactive templates refer to short phrases or behavioral patterns that are pre-designed and stored to attract or restore the audience's attention. These templates are not generated in real time by a large model, but are prepared offline and validated interactive corpora, such as "Baby, guess what will happen next?", "Listen carefully, the little rabbit is in big trouble!", or "Would you like to count how many apples are on the tree?". Each template is usually associated with a specific type of attention shift or the age group of the audience. For example, templates for visual shifts emphasize audio cues or suspense, while templates for drowsiness introduce action instructions ("Stretch, let's continue"). In the specific implementation, the executing entity can maintain a template table and select the most matching template based on the currently detected shift type, the age parameter in the audience identity setting, and the plot progression in the current story state object. The selection process can be a simple rule matching or can be achieved using a lightweight classifier. Once selected, the executing entity treats the template text as an instant streaming text segment and directly sends it to the speech synthesis module for playback, without going through a complex large model generation process, thus ensuring extremely low response latency.
[0067] In this embodiment, to avoid the template insertion interrupting the original story narrative, the current text generation task needs to be paused, the interactive template played, and the previous generation resumed after the playback was completed. This can be achieved using a priority queue, where the audio segment generated by the interactive template has a higher playback priority than ordinary story segments, allowing for immediate interruption of the current playback (if necessary) or insertion before the next segment. Alternatively, the template content can be naturally embedded into the next sentence of the story without interruption, for example, by replacing the original next sentence with the template content, or by appending interactive statements after the story sentences. Since the template text is usually very short (within a few dozen words), its intrusiveness to the overall narrative is low, yet it effectively attracts attention. This approach improves the ability to proactively maintain audience engagement, allowing for timely intervention when attention wanes without requiring user input, thereby enhancing the continuity of storytelling and the user experience.
[0068] In this embodiment, the executing entity generates streaming text segments based on preset interactive templates in the following manner: The audience's attention state is input into a large model to generate reasons for attention shift. The audience's attention state refers to a set of quantified features output in real time by an image or voice acquisition device, such as the audience's gaze direction, head posture, eye opening / closing degree, facial expression, or changes in tone, speech rate, and energy in the speech. Streaming text segments are generated based on interactive templates matched to the reasons for attention shift. The interactive templates are pre-designed short text libraries used to restore audience attention, and each interactive template is associated with several applicable tags for reasons for attention shift. For example, a template for the reason of "repetitive plot" could be "Baby, there's going to be a big twist! Guess who's coming?", and a template for the reason of "external interference" could be "Did a little bird fly by? Let's continue listening to what the little bird in the story is doing?". Specifically, before inputting the audience's attention state into the large model, it can first pass through a preprocessing module to transform it into a semantic description, such as "the audience's head deviates more than 30 degrees to the left for 3 seconds" or "the audience emits a long sigh." Preprocessing can be done using rule mapping or a small classifier, with the aim of converting multidimensional numerical features into natural language descriptions that the large model can understand. Then, the preprocessed audience attention state is concatenated with the plot context in the current story state object (such as the content summary of the current narrative node, recently played text fragments) to construct a cue which is then fed into the large model. This cue requires the large model to infer possible reasons for the attention shift, such as "Is the story too slow?", "Is the current plot making the audience feel scared or bored?", or "Is the audience being disturbed by the external environment?" Based on its common-sense reasoning ability regarding human attention psychology, the large model outputs one or more possible reasons for the shift, in the format of natural language sentences or structured labels, such as "the audience feels the plot is repetitive and lacks originality" or "a sudden external sound attracted the audience's attention." This output is called the attention shift reason. After identifying the causes of attention shifts, the executing agent matches the causes of attention shifts output by the large model with the labels of each template based on keyword or semantic similarity calculations, selecting the template with the highest similarity. If the large model outputs multiple causes, the most relevant template can be selected by weighted confidence. After selecting an interactive template, the executing agent outputs the template text as a streaming text segment. Compared to the method of randomly or polling directly from a fixed template library, this embodiment uses the large model to infer the causes of the shifts before matching, making the generated interactive content more causal and explanatory and personalized for the audience's current attention problems. For example, not all attention shifts are addressed with the same phrase "Listen carefully!", but different strategies such as "Let me explain" or "Let's play a fast-forward game" are used depending on the cause, such as "The story is too difficult to understand" or "The content is too simple and boring".This cause-matching mechanism significantly improves the relevance and effectiveness of interactive templates, enabling listeners to feel the speaker's understanding and care while actively attracting their attention.
[0069] The audience state monitoring method disclosed in this embodiment continuously monitors the audience's state using image acquisition devices and / or voice acquisition devices, capturing changes in audience attention in real time. Once an attention shift is detected, a streaming text segment is generated based on a preset interactive template. This effectively draws the audience back into the story content, extending the duration of the interaction and enhancing immersion. The storytelling process is no longer a passive, one-way output, but exhibits intelligent characteristics similar to a human storyteller, paying attention to audience reactions and making adjustments accordingly. This significantly improves the user experience and educational value of the audio story playback process. Simultaneously, the introduction of interactive templates avoids the computational overhead of generating complex interactive content in real time, ensuring fast response times.
[0070] Based on any of the above embodiments, in order to demonstrate the emotional care capabilities of the companion-type intelligent agent, this embodiment further discloses the following technical solution: When a listener's state is detected to be generating negative emotions, the currently played streaming text segment is saved as a temporary text segment. Negative emotions refer to non-positive emotional states identified by image acquisition devices (such as facial expression recognition detecting fear, sadness, or anger) or voice acquisition devices (such as voice emotion recognition detecting crying or anxious tones). The temporary text segment records the last complete narrative content heard by the listener before the emotional change, typically including the most recently played sentence or plot unit. Its purpose is to preserve the narrative context before the negative emotions were triggered when adjusting the story content later, so that a smooth continuation can be achieved after soothing, or to analyze the cause of the emotions. Then, a large model combined with the current story context is used to analyze the negative emotions, determining the type of negative emotion, such as fear, sadness, anger, or frustration. Furthermore, an estimate of the emotion intensity (such as mild, moderate, or severe) and possible... The triggering cause (e.g., "because the villain's behavior in the story is too frightening") provides a basis for selecting subsequent soothing strategies. Next, based on the type of negative emotion, a matching preset soothing strategy is determined. These preset soothing strategies are offline-designed narrative adjustment schemes for different emotion types. For example, for the emotion of "fear", strategies may include "pausing the thrilling plot and switching to a warm daily scene", "introducing a friendly supporting character to comfort the protagonist", or "reducing the tension of the sound effects". For the emotion of "sadness", strategies may include "having the positive character in the story say encouraging words", "speeding up the plot to a positive ending", or "inserting a heartwarming interlude". Each strategy comes with specific state modification instructions, such as modifying plot progress markers, adjusting character emotional states, or changing tone parameters. The executing entity selects the most appropriate soothing strategy based on the parsed emotion type (combining intensity if necessary). Finally, the story state object is adjusted based on temporary text fragments and preset soothing strategies, and a streaming text fragment is generated based on the adjusted story state object and identity setting values. Adjustments may include reverting narrative nodes to the positions corresponding to temporary text fragments to avoid continuing to play content that triggers negative emotions; modifying plot progression markers to skip upcoming high-risk scenes or jump directly to pre-set safe scenarios; updating the emotion fields in the character relationship status table (e.g., reducing the "threat level" of villainous characters); or adding a "soothing mode" marker to proactively add comforting dialogue or calming descriptions during subsequent generation. The adjusted story state object is then recombined with the originally locked identity settings (narrator's voice, self-identification, etc. remain unchanged), triggering the resumption of streaming text fragment generation. The generated text fragments no longer follow the original plotline but instead adhere to soothing strategies, producing content that alleviates negative emotions in the audience.For example, the original story might be about to describe the appearance of a monster, but the revised version would be, "Don't worry, it's just the shadow of the curtains being blown by the wind. The little bear bravely went over to take a look." This approach allows for dynamic modification of the narrative direction while protecting the listener's emotional experience, enhancing the emotional care capabilities of the companion-like intelligent agent, and avoiding the user experience disruption caused by abruptly stopping or skipping the story.
[0071] Building upon the aforementioned embodiments, to provide a companion-style storytelling experience while adding a remote safety barrier for young or sensitive users, this embodiment further discloses the following technical solution: When negative emotions are detected to persist for more than a preset time threshold, a notification is sent to the emergency contact based on pre-bound emergency contact information. Here, "negative emotions persisting for more than a preset time threshold" means that the system's determination of the listener's emotional state is not based on instantaneous facial expressions or vocal characteristics, but rather requires that the same type of negative emotion (such as fear, sadness, or anger) remain stable within a continuous time window, and that the length of this window exceeds a pre-configured value (e.g., 30 seconds, 1 minute, or adjustable according to the listener's age). This design avoids triggering unnecessary notifications due to brief emotional fluctuations in the listener (such as being startled by a sudden noise and then quickly recovering); only truly persistent negative emotions are considered abnormal situations requiring external intervention. The executing entity can maintain an emotion state timer. Whenever the emotion recognition module outputs a negative emotion label, if the label is the same as the previous output (or within the tolerance error), the timer is incremented. Once the timer reaches a threshold, this step is triggered. If the emotion changes to neutral or positive during the process, the timer is reset. Pre-bound emergency contact information refers to contact data pre-entered by parents or guardians during the initial device configuration or user setup phase. This typically includes the contact's name, relationship (e.g., "Mom," "Dad"), contact information (phone number, instant messaging account, application user ID), and optional emergency notification priority. This information is encrypted and stored in the device's local secure storage or cloud user account and can be updated at runtime through privacy permission management. The executing entity can prioritize sending an SMS to the bound phone number via the network, containing the message, "Your child continues to show fear / sadness while using the story device; please pay attention." If the device supports cellular or internet connectivity, notifications can also be pushed through the accompanying parent app, accompanied by richer information such as emotion type, duration, and text or screenshots of the story segment that triggered the emotion (if the camera is available and the privacy policy allows). Through the above methods, this embodiment enables parents to be promptly informed and take appropriate measures when their child experiences persistent emotional distress, significantly improving the product's safety assurance capabilities.
[0072] Based on any of the above embodiments, to further clarify the generation method of user timbre information, this embodiment further discloses the following technical solution: The executing entity continuously collects audio data during user interaction and performs validity detection on the audio data; when the cumulative result of valid audio segments reaches a preset timbre modeling threshold, a user timbre model and corresponding timbre information are generated based on the valid audio segments. The user interaction process refers to the user's natural behavior when using a voice terminal daily, such as issuing voice commands to the device (e.g., "How's the weather today?", "Tell me a story"), answering questions from the device, or engaging in casual conversation with the device. This audio data is collected in the background without the user's awareness, rather than being triggered by a dedicated "start recording" function. Since the collected raw audio usually contains interference components such as environmental noise, silent segments, echoes from the device's own playback, and non-user voices, it is necessary to perform validity detection on the audio data and retain valid audio segments that meet preset quality conditions. Validity detection refers to quality screening of each collected audio segment (e.g., a single sentence segmented by voice activity detection) and retaining those segments that meet preset quality conditions as valid audio segments. Preset quality conditions can be defined from multiple dimensions: the signal-to-noise ratio should be higher than a threshold (e.g., 15 dB) to ensure clear speech; the segment length should be within a certain range (e.g., 1 to 10 seconds) to avoid being too short to extract stable features or too long to contain excessive redundancy; the audio should not exhibit clipping or significant distortion; and voiceprint recognition or speaker confirmation is required to determine whether the segment belongs to the target user (rather than other family members or visitors). Only audio segments that simultaneously meet these conditions will be retained and accumulated. The significance of validity detection lies in eliminating low-quality or irrelevant audio data, ensuring that the samples subsequently used for timbre modeling are pure and informative. The timbre modeling threshold can be set according to the total duration of valid audio, for example, accumulating 5 or 10 minutes of valid speech. Since modern timbre cloning (or speaker adaptation) technology no longer requires tens of minutes of high-quality recordings, a few minutes of everyday conversation segments, after augmentation and filtering, are sufficient to fine-tune a usable personalized timbre model. When the accumulated result reaches the timbre modeling threshold, the executing entity automatically triggers the timbre model generation process. This process is typically based on a deep learning architecture: starting with a basic multi-speaker or general timbre model, the model is fine-tuned or trained using accumulated effective audio clips, shifting the spectral characteristics of the model's output sound towards the target user's timbre. After training, a user-specific timbre model is output, and a unique identifier for the timbre information is assigned to it. This identifier can be a model index, an embedding vector, or a storage path for a set of model parameters, stored in the identity settings for subsequent speech synthesis. This approach significantly lowers the barrier to entry for personalized voice storytelling, allowing ordinary family users to seamlessly acquire a personalized timbre, thereby enhancing the naturalness and intimacy of the companionship experience.
[0073] To enhance understanding, this disclosure also provides a specific implementation scheme in conjunction with a particular application scenario, such as... Figure 7a and Figure 7b As shown.
[0074] Family scene: In the evening, the mother and her 5-year-old twin children (brother Xiao Le and sister Xiao Ya) are sitting on the sofa in the living room, with the smart speaker on the coffee table.
[0075] Initial voice data collection: Over the past week, the mother used the speaker daily to check the weather and set alarms. The speaker's backend implicitly collected approximately 6 minutes of effective voice clips from the mother. Once the modeling threshold was reached, a personalized voice model for the mother (voice ID: `mom_voice_01`) was automatically generated. The mother was completely unaware of this.
[0076] Step 1: Voice Input. Mom says to the speaker, "Little D, Little D, tell my two babies a story about a kitten fishing using my voice."
[0077] Step 2: Identity Resolution. The speaker converts speech to text, performs semantic analysis, and extracts: Narrator's identity parameters: "I" (confirmed as mother based on voiceprint). Listener identity parameter: "Two Babies" (mapped to Xiao Le and Xiao Ya) Step 3: Identity Locking. The system establishes the following identity settings: the speaker refers to herself as "Mom," addresses the audience as "babies," uses a gentle and patient tone, and has her voice information set to `mom_voice_01`. This identity is locked in the current session until Mom explicitly requests a switch.
[0078] Step 4: Timbre Data Sufficiency Determination. The system determines whether a timbre model exists.
[0079] Step 5: Voice Data Acquisition. The system check found that the voice model (`mom_voice_01`) did not exist, so it continued to collect valid audio data of the mother (voice ID: `mom_voice_01`).
[0080] Step 6: Story State Object Construction and Initialization. The system checks and finds that the voice model already exists (`mom_voice_01`), meeting the requirements. It skips to continue data collection and proceeds to construct and initialize the story state object. The system automatically determines the following based on the identity settings and the "Cat Fishing" theme: Story type: Fable + everyday anecdotes Initial scene: By the river, under bright sunshine Characters: Mother Cat, Kitten (Xiao Le's character), White Kitten (Xiao Ya's character) Narrative difficulty: Low (suitable for ages 5, short sentences, repetitive dialogue) Initialize the story state object, which includes the narrative node (opening), the character relationship table (the mother cat teaches the kittens to fish), and the plot marker (beginning).
[0081] Step 7: Streaming Text Generation. The large model generates the first text segment using the mother's language style and the current state object: "My darlings, today Mommy will tell you the story of the kitten fishing. The mother cat brought her kittens to the riverbank."
[0082] Step 8: Streaming speech synthesis and playback. The above text is synthesized using the mother's voice `mom_voice_01` and played through the speaker. The two children listen quietly.
[0083] Step 9: User Interruption Detection (Parallel Monitoring). While the speaker plays, the microphone continuously monitors the environment. When the audio reaches the line "Mommy Cat says: 'Focus on fishing...'", Xiaoya suddenly asks, "Mommy, will the kitten catch the butterfly?" The system detects non-story-related voice input and determines that an interruption event has occurred.
[0084] Step 10: Was an interruption detected? If no interruption is detected, the system continues to generate streaming text.
[0085] Step 11: State Update and Restart. When the system detects an interruption, it updates the story state object, regenerates the streaming text fragment, and plays it.
[0086] Save the text clip that has already been played: "Mom Cat said: 'Concentrate on fishing...'" (Only save the part that was played in its entirety).
[0087] Convert Xiaoya's question to text and input it into the large model to analyze the intent: Intent type = "Question", Focus object = "Kitten catching butterflies", Time anchor point = "Current plot".
[0088] Update story status object: Add "Butterfly" as a temporary supporting character to the character relationship table, and revert the narrative node to the position "after the cat mother speaks".
[0089] Based on the updated state object and the original identity settings, the following streaming text is regenerated: "Good question, baby. The kitten saw a butterfly fly by and was about to chase it when its mother gently said, 'If you chase the butterfly, your fishing rod will be dragged away.' The kitten held back and continued to stare at the water."
[0090] The speaker continued playing Mom's voice. Xiaoya received a satisfactory answer, and Xiaole listened even more attentively.
[0091] After that, the story was interrupted twice more—once when Xiao Le asked to "tell the cat mom what she said again" (triggering the repetition mechanism, the system rolled back the node and re-outputted), and once when both children laughed at the same time, causing their attention to be distracted (the speaker detected the shift in gaze through the camera and inserted an interactive template: "Babies, guess if the float moved?"). Each interruption followed the... Figure 1 The cycle continues, maintaining the mother's identity and voice throughout. After the story ends, the system saves the final story state object for easy continuation next time.
[0092] This implementation extracts narrator and listener identity parameters from user input commands and constructs structured identity settings. These identity settings are used as implicit information in the calculation of each streaming text segment, avoiding the character drift problem common in long narratives. This ensures that the narrator's self-identification, tone, and emotional inclination remain highly consistent throughout the story playback, providing users with a stable and reliable companionship experience. Then, the story state object initialized based on the identity settings continuously records the current story progress, character relationships, and plot nodes, so that the generation of subsequent text segments is always driven by the narrative state rather than relying on a flat historical record. This improves the logical coherence of generating long or continuous stories and prevents plot breaks or inconsistencies. Finally, by using a collaborative approach of streaming text generation and streaming speech synthesis, the first text segment can be synthesized and played immediately after its generation according to the pre-associated timbre information in the identity settings, without waiting for the complete story content to be prepared. This significantly reduces the user's perceived initial response delay, making the acoustic characteristics of the output voice firmly bound to the narrator's identity. Even if the story is long or involves multiple rounds of interaction, the listener can still obtain a consistent personalized timbre experience. This embodiment focuses on identity locking and state-driven operation, enabling voice interaction devices to have low latency, high coherence, and identity-consistent real-time voice storytelling capabilities. It achieves real-time voice story playback with identity locking, state-driven operation, and interruptible resume capability, significantly improving the user experience in smart companionship scenarios.
[0093] Further reference Figure 8 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a voice story playback device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0094] like Figure 8As shown, the audio story playback device 800 of this embodiment may include: a parameter extraction unit 801, an identity construction unit 802, an object initialization unit 803, and a voice generation unit 804. The parameter extraction unit 801 is configured to extract narrator identity parameters and listener identity parameters from the storytelling instruction input by the user; the identity construction unit 802 is configured to construct identity setting values based on the narrator identity parameters and listener identity parameters; the object initialization unit 803 is configured to initialize the story state object based on the identity setting values and generate streaming text segments based on the identity setting values and the real-time story state object; wherein the identity setting values are used as implicit information in the generation of each text segment; the voice generation unit 804 is configured to generate audio story segments from each sequentially generated streaming text segment according to the timbre information in the identity setting values and play them to the target listener; wherein the target listener is determined based on the listener identity parameters.
[0095] In this embodiment, the specific processing of the parameter extraction unit 801, identity construction unit 802, object initialization unit 803, and voice generation unit 804 in the voice story playback device 800, and the resulting technical effects, can be found in reference to [reference needed]. Figure 2 The relevant descriptions of steps 201-204 in the corresponding embodiments will not be repeated here.
[0096] In some optional implementations of this embodiment, the parameter extraction unit 801 is further configured as follows: an instruction conversion module is configured to convert the user-inputted storytelling instruction into storytelling text; and a parameter extraction module is configured to extract narrator identity parameters and audience identity parameters from the storytelling text.
[0097] In some optional implementations of this embodiment, the object initialization unit 803 is further configured as follows: a first initial object generation module is configured to determine the story type, initial scene, story characters and narrative difficulty according to the identity setting value, and generate an initial story state object according to the story type, initial scene, story characters and narrative difficulty; a first style and material determination module is configured to determine the language style corresponding to the identity setting value and the story material corresponding to the current story state object using a preset large model; and a first text generation module is configured to generate streaming text fragments based on the language style and story material.
[0098] In some optional implementations of this embodiment, the object initialization unit 803 is further configured as follows: an interruption text saving module, configured to save the currently played streaming text segment as an intermediate text segment in response to detecting interruption information from the target audience during playback; an interruption semantic generation module, configured to use a large model to perform intent parsing on the interruption information and generate interruption semantic information; and an object updating module, configured to update the current story state object based on the intermediate text segment and the interruption semantic information, and generate a streaming text segment based on the updated story state object and the identity setting value.
[0099] In some optional implementations of this embodiment, the identity construction unit 802 is further configured as: a multi-identity construction module, configured to extract at least two narrator identity parameters from the storytelling instruction, and construct multi-identity setting values based on the at least two narrator identity parameters and the listener identity parameters, wherein the narrator identity parameters include a character identity identifier and the character timbre information corresponding to the character identity identifier; and the object initialization unit 803 is further configured as: a second initial object generation module, configured to determine the story type, initial scene, story characters, and narrative difficulty based on the multi-identity setting values, and generate an initial story state object based on the story type, initial scene, story characters, and narrative difficulty; a second style and material determination module, configured to determine the language style corresponding to the multi-identity setting values and the story material corresponding to the current story state object using a preset large model; and a second text generation module, configured to generate streaming text fragments based on the language style and story material.
[0100] In some optional implementations of this embodiment, the object initialization unit 803 is further configured as follows: an audience monitoring module, configured to monitor the audience state using an image acquisition device and / or a voice acquisition device; and an interactive text generation module, configured to generate streaming text fragments according to a preset interactive template in response to the detection of an attention shift in the audience state.
[0101] In some optional implementations of this embodiment, the interactive text generation module in the object initialization unit 803 is further configured to: input the audience's attention state into the large model to generate the attention shift cause; and generate a streaming text fragment based on the interactive template matched by the attention shift cause.
[0102] In some optional implementations of this embodiment, the object initialization unit 803 is further configured as follows: a negative emotion monitoring module, configured to save the currently played streaming text segment as a temporary text segment in response to detecting negative emotions in the listener's state; an emotion type determination module, configured to analyze the negative emotions using a large model to determine the type of negative emotions; a soothing strategy determination module, configured to determine a matching preset soothing strategy based on the type of negative emotions; an object adjustment module, configured to adjust the story state object based on the temporary text segment and the preset soothing strategy, and generate a streaming text segment based on the adjusted story state object and the identity setting value; and an emergency notification module, configured to send a notification message to the emergency contact based on the pre-bound emergency contact information in response to detecting that the negative emotions continue to exceed a preset time threshold.
[0103] In some optional implementations of this embodiment, the timbre information in the voice story playback device 800 is generated in the following way: audio data during user interaction is collected, and the audio data is validated to retain valid audio segments that meet preset quality conditions; in response to the cumulative result of valid audio segments reaching a preset timbre modeling threshold, the user's timbre model and the timbre information corresponding to the timbre model are generated based on the valid audio segments.
[0104] This embodiment exists as a device embodiment corresponding to the above method embodiment. The voice story playback device provided in this embodiment extracts the identity parameters of the narrator and the audience from the user's input commands and constructs a structured identity setting value. This identity setting value participates in the calculation as implicit information during the generation of each streaming text segment, avoiding the character drift problem common in long narratives. It ensures that the narrator's self-reference, tone, and emotional inclination remain highly consistent throughout the story playback process, thereby providing users with a stable and reliable companionship experience. Then, the story state object initialized based on the identity setting value continuously records the current story progress, character relationships, and plot nodes, making... The generation of subsequent text fragments is always driven by the narrative state rather than relying on a flat historical record, improving the logical coherence of long or continuous stories and preventing plot breaks or inconsistencies. Finally, the collaborative approach of streaming text generation and streaming speech synthesis allows for immediate synthesis and playback of the first text fragment based on pre-associated timbre information in the identity settings, without waiting for the complete story content to be prepared. This significantly reduces the user's perceived initial response latency, ensuring a strong bond between the acoustic characteristics of the output voice and the narrator's identity. Even with long stories or multiple rounds of interaction, listeners receive a consistently personalized timbre experience. This embodiment, with identity locking and state-driven mechanisms at its core, enables voice interaction devices to possess low-latency, highly coherent, and identity-consistent real-time voice storytelling capabilities. It achieves real-time voice story playback with identity locking, state-driven mechanisms, and the ability to resume interrupted playback, significantly improving the user experience in intelligent companionship scenarios.
[0105] According to embodiments of the present disclosure, the present disclosure also provides an electronic device, the electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to implement the voice story playback method described in any of the above embodiments when executed.
[0106] According to embodiments of this disclosure, this disclosure also provides a readable storage medium storing computer instructions that enable a computer to implement the voice story playback method described in any of the above embodiments when executed.
[0107] According to embodiments of this disclosure, this disclosure also provides a computer program product that, when executed by a processor, can implement the voice story playback method described in any of the above embodiments.
[0108] Figure 9A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0109] like Figure 9 As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 902 or a computer program loaded from storage unit 908 into random access memory (RAM) 903. RAM 903 may also store various programs and data required for the operation of device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.
[0110] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0111] The computing unit 901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as the voice story playback method. For example, in some embodiments, the voice story playback method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the voice story playback method described above can be performed. Alternatively, in other embodiments, the computing unit 901 can be configured to perform the voice story playback method by any other suitable means (e.g., by means of firmware).
[0112] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0113] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0114] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0115] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0116] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0117] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.
[0118] According to the technical solution of this disclosure, by extracting the narrator and listener identity parameters from user input commands and constructing structured identity setting values, these identity setting values are used as implicit information in the calculation of each streaming text segment generation process. This avoids the character drift problem common in long narratives and ensures that the narrator's self-reference, tone, and emotional inclination remain highly consistent throughout the story playback, thereby providing users with a stable and reliable companionship experience. Then, the story state object initialized based on the identity setting values continuously records the current story progress, character relationships, and plot nodes, so that the generation of subsequent text segments is always influenced by the identity setting values. By driving narrative state rather than relying on a flat historical record, the logical coherence of generating long or continuous stories is improved, preventing plot breaks or inconsistencies. Finally, the collaborative approach of streaming text generation and streaming speech synthesis allows for immediate synthesis and playback of the first text fragment based on pre-associated timbre information set in the identity parameters, without waiting for the complete story content to be prepared. This significantly reduces the user's perceived initial response latency, ensuring a strong bond between the acoustic characteristics of the output voice and the narrator's identity. Even with long stories or multiple rounds of interaction, listeners receive a consistently personalized timbre experience. This embodiment, with identity locking and state-driven mechanisms at its core, enables voice interaction devices to possess low-latency, highly coherent, and identity-consistent real-time voice storytelling capabilities. It achieves real-time voice story playback with identity locking, state-driven mechanisms, and the ability to resume interrupted playback, significantly improving the user experience in intelligent companionship scenarios.
[0119] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0120] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for playing a voice story, comprising: Extract narrator and audience identity parameters from user-input storytelling commands; An identity setting value is constructed based on the narrator identity parameters and the audience identity parameters; The story state object is initialized according to the identity setting value, and a streaming text fragment is generated based on the identity setting value and the real-time story state object; wherein, the identity setting value is used as implicit information in the generation of each text fragment each time; Each sequentially generated streaming text segment will be used to generate a voice story segment according to the timbre information in the identity setting value and played to the target audience; wherein the target audience is determined based on the audience identity parameter.
2. The method according to claim 1, wherein, The extraction of narrator and listener identity parameters from user-inputted storytelling instructions includes: Transform user-input storytelling commands into storytelling text; Extract the narrator's identity parameters and the audience's identity parameters from the storytelling text.
3. The method according to claim 1, wherein, The initialization of the story state object based on the identity setting value, and the generation of streaming text fragments based on the identity setting value and the real-time story state object, includes: The story type, initial scene, story characters, and narrative difficulty are determined based on the identity setting values, and an initial story state object is generated based on the story type, initial scene, story characters, and narrative difficulty. The language style corresponding to the identity setting value and the story material corresponding to the current story state object are determined using a preset large model; Streaming text fragments are generated based on the language style and the story material.
4. The method according to claim 3, wherein, The generation of streaming text fragments based on the identity settings and real-time story state objects includes: In response to detecting an interruption message from the target audience during playback, the currently played streaming text segment is saved as an intermediate text segment; The interruption information is analyzed using the large model to generate interruption semantic information; The current story state object is updated based on the intermediate text fragment and the interruption semantic information, and a streaming text fragment is generated based on the updated story state object and the identity setting value.
5. The method according to claim 1, wherein, The step of constructing identity setting values based on the narrator identity parameters and the audience identity parameters includes: In response to extracting at least two narrator identity parameters from the storytelling instruction, a multi-identity setting value is constructed based on the at least two narrator identity parameters and the listener identity parameters, wherein the narrator identity parameters include a character identity identifier and the character timbre information corresponding to the character identity identifier; The initialization of the story state object based on the identity setting value, and the generation of streaming text fragments based on the identity setting value and the real-time story state object, include: The story type, initial scene, story characters, and narrative difficulty are determined based on the multi-identity setting values, and an initial story state object is generated based on the story type, initial scene, story characters, and narrative difficulty. The language style corresponding to the multi-identity settings and the story material corresponding to the current story state object are determined using a pre-defined large model; Streaming text fragments are generated based on the language style and the story material.
6. The method according to any one of claims 1-5, wherein, The generation of streaming text fragments based on the identity settings and real-time story state objects includes: Use image acquisition equipment and / or voice acquisition equipment to monitor the audience's state; In response to the detection of an attention shift in the audience's state, a streaming text segment is generated based on a preset interaction template.
7. The method according to claim 6, wherein, The step of generating streaming text fragments based on a preset interactive template includes: The audience's attention state is input into the large model to generate the reasons for attention shift; Streaming text fragments are generated based on the interactive template matched to the cause of the attention shift.
8. The method according to claim 6, wherein, The generation of streaming text fragments based on the identity settings and real-time story state objects includes: In response to detecting negative emotions in the listener's state, the currently played streaming text segment is saved as a temporary text segment; The negative emotions are analyzed using the large model to determine their type. Determine a matching preset soothing strategy based on the type of negative emotion; The story state object is adjusted based on the temporary text fragment and the preset soothing strategy, and a streaming text fragment is generated based on the adjusted story state object and the identity setting value.
9. The method according to claim 8, further comprising: In response to the detection that the negative emotion continues for more than a preset time threshold, a notification message is sent to the emergency contact based on the pre-bound emergency contact information.
10. The method according to claim 1, wherein, The timbre information is generated in the following way: Collect audio data during user interaction, perform validity checks on the audio data, and retain valid audio segments that meet preset quality conditions; In response to the cumulative result of the effective audio segments reaching a preset timbre modeling threshold, the user's timbre model and timbre information corresponding to the timbre model are generated based on the effective audio segments.
11. A voice story playback device, comprising: The parameter extraction unit is configured to extract narrator identity parameters and audience identity parameters from the user-inputted storytelling instructions. An identity construction unit is configured to construct identity setting values based on the narrator identity parameters and the audience identity parameters; The object initialization unit is configured to initialize the story state object according to the identity setting value, and generate streaming text fragments based on the identity setting value and the real-time story state object; wherein the identity setting value is used as implicit information in the generation of each text fragment each time. The speech generation unit is configured to generate speech story segments for each sequentially generated streaming text segment according to the timbre information in the identity setting value and play them to the target audience; wherein the target audience is determined based on the audience identity parameters.
12. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the voice story playback method according to any one of claims 1-10.
13. A non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the audio story playback method according to any one of claims 1-10.
14. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the voice story playback method according to any one of claims 1-10.