Personalized child story generation system based on artificial intelligence

By integrating speech recognition, multimodal data analysis and dynamic story generation technology in the children's story generation system, it solves the problem that existing equipment is difficult to generate story content that meets their needs in real time based on individual differences in children, and achieves efficient and personalized children's story generation, improving user experience and educational effects.

CN120067336AInactive Publication Date: 2025-05-30SHENZHEN XINLIKANG ELECTRONICS CO LTD

Patent Information

Application Number
CN202510548168.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing children's story generation equipment is difficult to generate story content that meets their psychological needs in real time based on children's language expression habits, age, gender characteristics and interest preferences, resulting in insufficient personalization and interactivity of user experience.

Method used

A personalized children's story generation system based on artificial intelligence is designed. Through speech recognition, multimodal data analysis and dynamic story generation technology, the voice recognition module, information extraction module, story generation module and interactive adjustment module are integrated to respond to children's input and feedback in real time.

Benefits of technology

It realizes a seamless transformation from children's voice to personalized story content, significantly improving the flexibility and adaptability of content generation, and providing a high-quality personalized educational and entertainment experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067336A_ABST
    Figure CN120067336A_ABST
Patent Text Reader

Abstract

The invention discloses a personalized child story generation system based on artificial intelligence, and the system comprises a voice recognition module which is used for obtaining the original audio of a child, and extracting a text feature sequence based on the original audio of the child; the information extraction module is used for extracting semantic features, gender features and interest preference features from the text feature sequence; the story generation module is used for generating a personalized story based on the semantic features, the gender features and the interest preference features; the interaction adjustment module is used for dynamically updating the input of the story generation module based on the personalized story and the child original audio obtained through real-time interaction; real-time voice input analysis, multi-modal feature fusion and dynamic story generation technologies are utilized, and efficient hardware support is combined, so that seamless conversion from child voice to personalized story content is realized; compared with a scheme depending on a static database in the prior art, the method has the advantage that the flexibility and adaptability of content generation are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a generation system, specifically a personalized children's story generation system based on artificial intelligence. Background Art

[0002] With the rapid evolution of artificial intelligence technology, especially the technological breakthroughs in the fields of speech recognition, natural language processing, and deep learning, the application potential of AI in children's education and entertainment scenarios has become increasingly prominent. However, existing children's story generation devices generally rely on static story databases, with their content presenting a single and fixed nature, lacking the ability to dynamically respond to the individual differences in children's needs. For example, existing technologies are difficult to generate story content that meets their psychological needs in real time according to children's language expression habits, age stages, gender characteristics, and interest preferences, resulting in insufficient personalization and interactivity of the user experience.

[0003] Although existing intelligent story devices have certain voice interaction functions, their cores are still limited to preset story templates or limited script libraries, making it difficult to achieve dynamic content adjustment based on children's real-time input. The rise of large language models (LLMs) provides a technical foundation for generating diverse and contextualized text content, while multimodal data analysis further enhances the ability to deeply understand user input. However, existing devices have not fully integrated these cutting-edge technologies and cannot achieve a real breakthrough in personalization in the field of children's education and entertainment. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention provides a personalized children's story generation system based on artificial intelligence, which integrates multimodal data analysis to solve the technical problems raised in the background art.

[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A personalized children's story generation system based on artificial intelligence, including the following steps: A speech recognition module for obtaining children's original audio and extracting a text feature sequence based on the children's original audio; An information extraction module for extracting semantic features, gender features, and interest preference features from the text feature sequence; A story generation module for generating personalized stories based on semantic features, gender features, and interest preference features; An interactive adjustment module for dynamically updating the input of the story generation module based on the personalized story and the children's original audio obtained from real-time interaction.

[0006] In some of these embodiments, the speech recognition module includes: A language preprocessing sub-module for converting children's original audio into MFCC features; A speech-to-text sub-module for converting MFCC features into a readable text feature sequence; An emotion recognition sub-module for identifying the emotional classification representation of children based on the text feature sequence.

[0007] In some embodiments, the information extraction module includes: A semantic extraction sub-module for extracting implicit feature representations related to age, gender, and interests from the text feature sequence; A gender inference sub-module for determining the gender classification representation of children based on the children's original audio and the implicit feature representation of gender; An interest preference extraction sub-module for determining the interest preference representation of children based on the implicit feature representation of interests; In some embodiments, the story generation module includes: An intention recognition sub-module for determining the intention representation of children based on the children's original audio and the children's emotional classification representation; A dynamic adjustment sub-module for dynamically updating the generated story based on the children's intention representation and the children's original audio obtained from real-time interaction; A feedback integration sub-module for collecting the children's emotional classification representation in real time and using the real-time collected children's emotional classification representation as the input of the children's intention representation.

[0008] In some embodiments, the interaction adjustment module includes: A multi-modal input support sub-module for obtaining input instructions of children's speech, text, and gestures; An emotion perception and regulation sub-module for obtaining the corresponding children's emotional classification representation based on the input instructions of children's speech, text, and gestures; An adaptive difficulty adjustment sub-module for dynamically adjusting the language complexity and plot depth of the story according to the children's age, original audio input, and interaction history.

[0009] In some embodiments, the speech recognition module is an end-to-end deep learning model based on the Transformer architecture, and realizes speech recognition through the multi-head self-attention mechanism.

[0010] In some embodiments, the information extraction module is based on a pre-trained BERT model and extracts implicit feature representations in the text through bidirectional semantic encoding.

[0011] In some of these embodiments, the story generation module is based on the GPT generative large language model and adjusts the model output to match the language style and cognitive level of children by fine-tuning on a children's language dataset; the fine-tuning process uses supervised learning, and the optimization objective is to minimize the cross-entropy loss between the generated text and the annotated children's stories.

[0012] In some of these embodiments, the interactive adjustment module uses a context memory mechanism to interact with historical data in real time through a recurrent neural network.

[0013] The present invention provides a personalized children's story generation system based on artificial intelligence, which has the following beneficial effects: The present invention utilizes real-time speech input parsing, multi-modal feature fusion, and dynamic story generation technologies, combined with efficient hardware support (such as a neural network processor NPU and a microphone array), to achieve a seamless conversion from children's speech to personalized story content. Compared with the existing technology that relies on a static database, the present invention significantly improves the flexibility and adaptability of content generation.

[0014] With the continuous development of artificial intelligence technology and the increasing growth of children's education needs, the present invention shows broad application potential in the fields of family education, school teaching, intelligent toys, etc.; by supporting multi-language expansion, emotion perception, and age-appropriate adjustment, this device can provide high-quality personalized education and entertainment experiences for children worldwide, further promoting the fairness and accessibility of educational resources.

[0015] The present invention not only improves the user experience and cognitive development of children, but also embeds educational content in an entertaining way, providing a new tool for parents and educational institutions. At the same time, its highly personalized and multi-sensory interaction characteristics endow the device with strong market competitiveness. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a structural block diagram of a personalized children's story generation system based on artificial intelligence according to the present invention; Figure 2 is a schematic diagram of the generation process of a personalized children's story generation system based on artificial intelligence according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0018] Refer toFigure 1 , the present invention provides a personalized children's story generation system based on artificial intelligence, and the generation system includes: A speech recognition module, configured to obtain the original audio of a child and extract a text feature sequence based on the original audio of the child; An information extraction module, configured to extract semantic features, gender features, and interest preference features from the text feature sequence; A story generation module, configured to generate a personalized story based on semantic features, gender features, and interest preference features; An interactive adjustment module, configured to dynamically update the input of the story generation module based on the personalized story and the original audio of the child obtained through real-time interaction.

[0019] Among them, the working process of the personalized children's story generation system is as follows:

[0020] 1. Voice input: The child inputs a voice signal through a microphone array.

[0021] 2. Signal processing: The speech recognition module decodes the signal into text and extracts emotional features at the same time.

[0022] 3. Information parsing: The information extraction module analyzes the text and speech features and outputs structured parameters (age, gender, preference, emotion).

[0023] 4. Story generation: The story generation module generates an initial story based on the parameters, and the NPU accelerates the inference process.

[0024] 5. Multimodal output: The display screen renders story illustrations, and the speaker plays the dubbed audio.

[0025] 6. Real-time adjustment: The child's feedback triggers the interactive adjustment module, and the model dynamically inserts new plots and updates the output.

[0026] Specifically, the present system adopts a modular design, and the modular design includes: A speech recognition module: configured to receive the voice input of a child, convert it into text, and extract emotional features.

[0027] An information extraction module: parse information such as the age, gender, and interest preference of a child from the speech text and emotional features.

[0028] A story generation module: generate personalized story content based on the parsed information based on a large language model.

[0029] An interactive adjustment module: support the child to intervene in the development of the story in real time through voice to achieve dynamic adjustment.

[0030] For the present system to achieve efficient operation and real-time interaction, its hardware configuration is as follows: Central Processing Unit (CPU): Adopts a multi-core processor, responsible for task scheduling and data processing.

[0031] Neural Network Processor (NPU): Specifically designed for artificial intelligence model inference, accelerating speech recognition and story generation.

[0032] Memory: Includes high-speed random access memory (RAM) and non-volatile storage units (such as NAND flash), used to store model parameters, user data, and temporary files.

[0033] Microphone array: With a multi-channel design, supporting far-field sound pickup and noise suppression, improving the quality of voice input.

[0034] In this embodiment, in order to implement the story generation described in the present invention, the modular design of the present invention includes: Speech recognition module The main functions of the speech recognition module include accurately converting the speech input of children into text (speech-to-text) and extracting emotional features from the speech and recognizing the emotional state (emotion recognition). In the speech-to-text sub-module, an end-to-end deep learning model based on the Transformer architecture is adopted, and high-precision speech transcription is achieved through the multi-head self-attention mechanism. In order to adapt to the characteristics of children's speech, the model is pre-trained on a children's speech dataset containing different ages, genders, and accents. This module first preprocesses the original audio signal, then extracts MFCC features, and inputs them into the Transformer model to generate a text sequence.

[0035] Among them, the speech recognition module has the following functions: Speech-to-text: It is implemented through the speech-to-text sub-module, specifically responsible for converting the original audio of children into a readable text sequence. Its core is an end-to-end speech recognition model based on the Transformer architecture. Due to its powerful sequence modeling ability and parallel computing advantages, the Transformer has gradually replaced traditional recurrent neural networks (RNNs) and convolutional neural networks (CNNs) in the field of speech recognition in recent years. Through the multi-head self-attention mechanism, this model can efficiently capture long-range dependencies in speech signals, especially suitable for dealing with non-standard pronunciations and speech rate fluctuations common in children's speech.

[0036] Among them, the model architecture is: The Transformer model consists of two main parts: an encoder and a decoder, including: Encoder: The encoder is stacked by 6 Transformer encoding layers, and each layer contains two core sub-layers: Multi-Head Self-Attention Sub-Layer: It calculates multiple attention distributions in parallel through the multi-head mechanism. Each attention head independently learns different subspace representations of speech features. The calculation of multi-head self-attention is based on Scaled Dot-Product Attention, and the formula is:

[0037] where, are the query, key, and value matrices respectively, is the dimension of the key (usually set to 64), and the scaling factor is used to alleviate the numerical instability of high-dimensional matrix products. The number of attention heads is set to 8, enhancing the feature expression ability of the model.

[0038] Feed-Forward Neural Network Sub-Layer: It adopts a point-wise feed-forward network, which includes two layers of linear transformations and a ReLU activation function. The formula is:

[0039] where, and are weight matrices, and are bias terms, and the hidden dimension is set to 2048.

[0040] The encoder converts the input speech feature sequence into a high-dimensional representation, retaining the temporal information and semantic features.

[0041] Decoder: The decoder is also composed of 6 Transformer decoding layers. Each layer contains three sub-layers: Masked Multi-Head Self-Attention Sub-Layer: When generating the text sequence, through the masking mechanism, it ensures that the current time step can only focus on the previous outputs, avoiding leakage of future information.

[0042] Encoder-Decoder Attention Sub-Layer: It fuses the output of the encoder through the cross-attention mechanism, enhancing the alignment ability between the speech signal and the text sequence.

[0043] Feed-Forward Neural Network Sub-Layer: Consistent with the encoder, it is used to further transform the features. The decoder gradually generates the text sequence and finally outputs a transcription result that highly matches the input speech.

[0044] Positional Encoding: Transformer itself does not have the ability to model time series. Therefore, positional encoding is generated through sine and cosine functions to inject time order information into each element in the input sequence. The formula is:

[0045] Among them, is the position index, is the dimension index, is the model dimension (set to 512).

[0046] Speech feature extraction: Before inputting into the Transformer, the speech signal needs to go through preprocessing and feature extraction. This module uses Mel Frequency Cepstral Coefficients (MFCC) as the main features, and its extraction process is as follows: Pre-emphasis: Enhance the high-frequency components through a first-order high-pass filter y(n)=x(n)−0.97x(n−1) to compensate for the oral radiation effect.

[0047] Framing: Segment the speech signal into short-time frames of 20 - 40 ms with a frame shift of 10 ms to ensure the signal's smoothness.

[0048] Windowing: Apply a Hamming window to each frame to reduce spectral leakage.

[0049] Fast Fourier Transform (FFT): Calculate the power spectrum of each frame with a length of 512 points.

[0050] Mel filter bank: Simulate the non-linear frequency perception of the human ear through 40 triangular filters and output the logarithmic energy spectrum.

[0051] Discrete Cosine Transform (DCT): Transform the logarithmic energy spectrum into 13-dimensional MFCC coefficients, and append the first-order and second-order differences to form a 39-dimensional feature vector.

[0052] Finally, the MFCC feature sequence is arranged in chronological order and fed into the Transformer encoder.

[0053] Exemplarily, the training and optimization of the model involved in the speech recognition module are as follows: To adapt to the characteristics of children's speech, the model is pre-trained on a children's speech dataset containing different ages, genders, and accents.

[0054] The following optimization strategies are adopted during the training process: Data augmentation: Enhance the model's robustness by adding background noise (such as white noise, ambient sound), speed variation (0.9x to 1.1x), and pitch variation (±2 semitones).

[0055] CTC loss function: Use the Connectionist Temporal Classification Loss to handle the problem of misalignment between the input speech and the output text length. The formula is: ; Among them, is the path mapping function, is the input feature sequence, is the target text sequence.

[0056] Beam search decoding: When inferring, the beam search algorithm with a beam width of 10 is used, and the text sequence generation is optimized by combining the language model scores.

[0057] Furthermore, the speech recognition module also has: Emotion recognition: It is implemented through the emotion recognition sub-module. Specifically, it aims to analyze the emotional state in children's speech (such as "happy", "calm", "angry"), and is implemented by combining the convolutional neural network (CNN) and the long short-term memory network (LSTM). This hybrid architecture makes full use of the local feature extraction ability of CNN and the temporal modeling ability of LSTM.

[0058] Feature extraction: Its emotion recognition depends on various speech features, including: MFCC: Shared with speech-to-text, capturing spectral characteristics.

[0059] Fundamental frequency (F0): Extracted by the autocorrelation method, reflecting pitch changes and serving as a key indicator for emotional expression.

[0060] Energy features: Including short-time energy and zero-crossing rate, reflecting loudness and intensity.

[0061] Speech speed: Calculate the rhythm change through pitch detection and frame duration. And the emotion recognition sub-module applies the trained emotion recognition model vehicle, and its architecture is as follows The emotion recognition model adopts a CNN-LSTM hybrid architecture, and the specific structure is as follows: CNN layer: It contains 3 convolutional layers. Each layer uses a 3x3 convolutional kernel, with a stride of 1, and the number of channels is 32, 64, and 128 in sequence. After each convolutional layer, there is a max pooling layer (2x2 pooling kernel, stride 2) to extract local emotion patterns. The ReLU activation function is used to enhance the non-linear expression ability.

[0062] LSTM layer: Receives the feature sequence output by CNN, and adopts a bidirectional LSTM structure with 256 hidden units.

[0063] , ,

[0064] Among them, is the input, is the hidden state at the previous moment, is the sigmoid function.

[0065] Attention mechanism: An attention layer is introduced after the LSTM to calculate the weights for each time step:

[0066] Among them, the speech recognition module maps the attention output to the emotion category space and uses the Softmax function to output the classification probability.

[0067] As the data input and output of the above model, its training process is as follows: The model is trained on a children's speech dataset with labeled emotion tags, and the optimization strategies include: Data balancing: Sampling minority-class emotion samples through the SMOTE algorithm.

[0068] Transfer learning: Based on an adult emotion pre-trained model, adapting to children's characteristics through fine-tuning.

[0069] Loss function: Using cross-entropy loss and adding L2 regularization to prevent overfitting.

[0070] Furthermore, for the information extraction module, its main task is to extract key information from children's speech transcription texts and related features, including age, gender, and interest preferences. This information provides an important basis for subsequent personalized story generation. This module comprehensively applies deep learning, natural language processing, and signal processing technologies, and is specifically divided into three sub-modules: semantic extraction, gender inference, and interest preference extraction. The technical implementation of each sub-module will be introduced in detail below.

[0071] The semantic extraction sub-module is constructed based on the pre-trained BERT (Bidirectional Encoder Representations from Transformers) model. Through its bidirectional Transformer architecture, BERT can effectively capture the context information in the text and is very suitable for processing the complex semantics in children's natural language expressions. Through multi-layer encoding, BERT generates context-related representations for each input word, thereby extracting implicit information related to age, gender, and interests.

[0072] In specific applications, the input speech transcription text is first tokenized and converted into a token sequence that can be processed by BERT, and then encoded through the model's multi-layer self-attention mechanism. Finally, the encoded representations are fed into the classification layer to predict the age group of children (such as toddlers, preschoolers, etc.) and interest labels (such as animals, adventures, etc.). To adapt to the characteristics of children's language, the pre-trained BERT model will be fine-tuned on a specific dataset that contains a large number of children's text samples labeled with age and interests. BERT's self-attention mechanism is its key component, and the calculation formula is:

[0073] Among them, are the query, key, and value matrices respectively, is the dimension of the key, used to scale the dot product to avoid overly large numerical values.

[0074] The gender inference sub-module combines voice features and text semantic information to achieve more accurate gender classification. Voice features mainly rely on the fundamental frequency (F0), which is the pitch of the voice signal and reflects the physiological characteristics of the speaker. Generally, there are differences in the fundamental frequency ranges of men and women. The text semantic information is extracted through the BERT model to capture the expression patterns (such as word usage habits) in children's language that may be related to gender.

[0075] In the processing flow, first, the fundamental frequency features are extracted from the voice signal, and then these features are fused with the text representations generated by BERT. The fused feature vectors are input into a support vector machine (SVM) classifier. By training the SVM model to find the optimal hyperplane in the feature space that distinguishes between men and women, gender prediction is completed. This gender prediction step includes: Fundamental frequency extraction: The fundamental frequency is calculated by the autocorrelation method, and the formula is:

[0076] Among them, is the voice signal, is the time delay, is the signal length.

[0077] SVM classification: The decision function of SVM is: ; Among them, is the weight vector, is the bias, is the input feature vector.

[0078] The interest preference extraction sub-module aims to identify children's interest tendencies from their texts to provide personalized content support for story generation. This sub-module combines two methods: keyword co-occurrence analysis and latent Dirichlet allocation (LDA) topic modeling. Keyword co-occurrence analysis constructs a co-occurrence matrix by counting the co-occurrence frequencies of interest-related words (such as "dinosaur" and "princess") in the text to initially identify children's interest points. LDA topic modeling goes further by inferring the latent topic distribution of the text to generate a probabilistic interest representation. For example, LDA may identify that a certain text contains two topics, "animals" and "adventure", and assign probabilities to each topic. Finally, by combining the results of co-occurrence analysis and LDA, a comprehensive interest vector is generated for subsequent story content planning.

[0079] Among them, the topic generation process of LDA is based on the following probability distributions:

[0080] Among them, is the topic distribution of the document, is the topic, is the word, and are hyperparameters.

[0081] For the story interaction module, which is an important part of the system, its goal is to enhance the interactivity between children and the generated story, making the experience more immersive and personalized. This module not only allows children to passively receive the story content, but also supports their real-time intervention and feedback during the story generation process, thus affecting the development direction of the story. The story interaction module mainly consists of three sub-modules: intention recognition, dynamic adjustment, and feedback integration. The following will elaborate on the technical implementation of each sub-module in detail.

[0082] The intention recognition sub-module is responsible for analyzing the speech or text input by children during the interaction, identifying their intentions and converting them into actionable instructions. For example, a child may say "Make the protagonist become a dragon" or "I want a happy ending", and the system needs to accurately understand these requests. This sub-module uses natural language processing (NLP) techniques and combines pre-trained intention classification models (such as BERT or RoBERTa) to perform semantic parsing on the children's input.

[0083] In specific operations, the system first converts the children's input into an embedding vector, and then uses a classifier to determine the intention category (such as "change character", "adjust plot", "add element"). To adapt to the diversity and non-standardization of children's language, the model will incorporate children's corpus data during training and use few-shot learning techniques to improve the generalization ability for new intentions.

[0084] Among them, the classification probability of intention recognition can be expressed as:

[0085] Among them: input is the children's input; embedding(input) is the vector representation of the input; and are the parameters of the classifier.

[0086] The dynamic adjustment sub-module modifies the story content being generated in real time according to the results of intention recognition, ensuring that the story can reflect the wishes of children. This sub-module is based on conditional generation technology and adjusts the generation direction of the language model by updating the input conditional vector. For example, if a child requests "make the story more adventurous", the system will add adventure-related keywords (such as "treasure", "forest") to the conditional vector.

[0087] To maintain the coherence of the story, the dynamic adjustment adopts an incremental generation strategy, that is, gradually expanding on the basis of the existing story fragments instead of regenerating the entire story. The specific methods include beam search combined with context and dynamic weight adjustment to ensure that the newly generated content seamlessly connects with the previous plot.

[0088] Among them, the generation probability of dynamic adjustment is:

[0089] Among them, is the generated story fragment, is the conditional input of the child's intention.

[0090] The feedback integration sub-module is responsible for collecting the emotional feedback (such as "like", "dislike") and explicit evaluations of children in the story interaction and using them to optimize subsequent story generation. This sub-module combines sentiment analysis and reinforcement learning techniques to adjust the generation strategy according to the children's responses.

[0091] Specifically, the system analyzes the children's speech intonation or text feedback through a sentiment classification model to judge their satisfaction; at the same time, the reinforcement learning algorithm uses the positive feedback of children as a reward signal to optimize the parameters of the language model. For example, if a child shows interest in the "magic" element, the system will increase the generation probability of related content.

[0092] To ensure real-time performance, the feedback integration adopts an online learning method to dynamically update the model weights without the need for offline retraining.

[0093] Among them, the reward function of reinforcement learning can be expressed as:

[0094] Among them, is the immediate reward at the t-th step (such as children's satisfaction), is the discount factor, is the feedback of children.

[0095] For the interactive adjustment module, which is the core part of the children's story generation system, it provides a personalized story experience for children through real-time interaction and dynamic adjustment. This module consists of three sub-modules: a reinforcement learning framework, context memory, and dialogue state tracking. Their functions and technical implementations were introduced in detail in the previous paragraph. Next, we will explore the further optimization and extended functions of this module to enhance the system's adaptability and user experience.

[0096] To further enrich the interaction methods, a multi-modal input support sub-module is introduced, allowing children to express feedback and instructions in various forms such as speech, text, and even gestures. This sub-module uses multi-modal fusion technology to integrate different input signals into a unified semantic representation. Specifically, speech input is converted into text through automatic speech recognition (ASR), text input directly undergoes natural language processing, and gesture input is parsed into semantic instructions through computer vision technology (such as convolutional neural network, CNN).

[0097] The system uses a Transformer-based fusion model to map multi-modal features into a shared representation space and dynamically adjusts the weights of each modality through the self-attention mechanism. For example, when a child says "Make the princess be braver" and makes a fist gesture at the same time, the system will synthesize the meanings of both, strengthening the manifestation of the "brave" attribute in the story.

[0098] The output representation of multi-modal fusion is:

[0099] Where: 、 、 represent the feature vectors of speech, text, and gestures respectively; 、 、 are the linear transformation matrices of the corresponding modalities; is the fused semantic representation.

[0100] The emotion perception and regulation sub-module analyzes the child's speech intonation and feedback content to identify their emotional state (such as excitement, disappointment), and adjusts the emotional tone of the story accordingly. The system uses an emotion classification model (such as a fine-tuned model based on BERT) to process text and speech features, and combines acoustic features (such as pitch, speech rate) for multi-modal emotion analysis.

[0101] In terms of regulation, the system dynamically adjusts the generation strategy according to the emotional state. For example, if it detects that the child is disappointed, the system may increase humorous elements or introduce characters the child likes. To achieve real-time performance, this sub-module uses a lightweight model and continuously optimizes the accuracy of emotion recognition through online learning.

[0102] The softmax output of emotion classification is:

[0103] Wherein: is the emotion category (such as positive, negative); is the input feature vector; 、 are the weights and biases of the emotion category; is the emotion probability.

[0104] The adaptive difficulty adjustment sub-module dynamically adjusts the language complexity and plot depth of the story according to the child's age, language ability and interaction history. The system uses a Bayesian inference model to evaluate the child's comprehension level and updates the user profile in combination with interaction data (such as the complexity of instructions).

[0105] For example, for younger children, the system tends to generate simple sentences and linear plots; for older children, more branching options and complex vocabulary are introduced. To ensure the smoothness of the adjustment, the system uses a smoothing factor to control the amplitude of the difficulty change and avoid abrupt experience interruptions.

[0106] The Bayesian update formula is:

[0107] Wherein: is the child ability parameter, is the interaction data; is the posterior probability; is the likelihood function; is the prior probability; is the normalization constant.

[0108] It should be particularly noted that the protection scope of the present invention is defined by the appended claims and their equivalent scope. Any improvement, adjustment or equivalent replacement based on the technical solution of the present invention shall be regarded as a natural extension of the present invention and fall within the protection scope of the present invention. The present invention not only achieves a revolutionary breakthrough in technology in the field of children's content generation, but also has far-reaching significance in promoting children's education innovation and enhancing social well-being.

[0109] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means.

[0110] The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more collections of available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVD ), or semiconductor media. The semiconductor media can be a solid-state drive.

[0111] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0112] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application.

Claims

1. A personalized children's story generation system based on artificial intelligence, characterized in that: include: A speech recognition module, used to obtain the original audio of the child and extract a text feature sequence based on the original audio of the child; An information extraction module, used to extract semantic features, gender features and interest preference features from the text feature sequence; Story generation module, used to generate personalized stories based on semantic features, gender features, and interest preference features; The interactive adjustment module is used to dynamically update the input of the story generation module based on the personalized story and the original audio of the child obtained through real-time interaction.

2. The artificial intelligence-based personalized children's story generation system according to claim 1, characterized in that: The speech recognition module comprises: The language preprocessing submodule is used to convert children’s raw audio into MFCC features; The speech-to-text submodule is used to convert MFCC features into readable text feature sequences; The emotion recognition submodule is used to identify children’s emotion classification representation based on text feature sequences.

3. The personalized children's story generation system based on artificial intelligence according to claim 1, characterized in that: The information extraction module includes: Semantic extraction submodule, used to extract implicit feature representations related to age, gender and interests from text feature sequences; A gender inference submodule is used to determine the child’s gender classification representation based on the acoustic features of the child’s original audio and the semantic representation of the text feature sequence; The interest preference extraction submodule is used to determine the children's interest preference representation based on the implicit feature representation of the interest.

4. The artificial intelligence-based personalized children's story generation system according to claim 1, characterized in that: The story generation module includes: The intention recognition submodule is used to determine the child's intention representation based on the child's original audio and the child's emotion classification representation; A dynamic adjustment submodule is used to dynamically update the generated story based on the child's intention expression and the child's original audio obtained through real-time interaction; The feedback integration submodule is used to collect children’s emotion classification representations and explicit evaluations in real time and use them to optimize the story generation process.

5. The personalized children's story generation system based on artificial intelligence according to claim 1 is characterized in that: The interactive adjustment module includes: Multimodal input support submodule, used to obtain children's voice, text and gesture input instructions; The emotion perception and regulation submodule is used to obtain the corresponding child emotion classification representation based on the input instructions of children's voice, text and gestures; An adaptive difficulty adjustment submodule is used to dynamically adjust the language complexity and plot depth of the story based on the child’s age, original audio input, and interaction history.

6. The artificial intelligence-based personalized children's story generation system according to claim 1, characterized in that: The speech recognition module is based on an end-to-end deep learning model of the Transformer architecture and realizes speech recognition through a multi-head self-attention mechanism.

7. The artificial intelligence-based personalized children's story generation system according to claim 1, characterized in that: The information extraction module is based on the pre-trained BERT model and extracts implicit feature representations in the text through bidirectional semantic encoding.

8. The artificial intelligence-based personalized children's story generation system according to claim 1, characterized in that: The story generation module uses GPT's generative large language model to fine-tune children's language style and cognitive level.

9. The artificial intelligence-based personalized children's story generation system according to claim 1, characterized in that: The interactive adjustment module adopts a context memory mechanism and uses a recurrent neural network to interact with historical data in real time.

Citation Information

Patent Citations

  • Sound-based method and apparatus for generating AR content and storage medium

    CN109065055A

  • Method and system for automatically generating story video in meta universe

    CN117177003A

  • Story video generation method and device, storage medium and equipment

    CN117332118A

Cited By

  • Story interactive active questioning method and system based on cognitive characteristics of children

    CN121119146A