Multi-modal picture book material fusion and dynamic presentation method based on AIGC

Through AIGC technology, picture texts, illustrations and audio are generated, and combined with AR technology, the problems of high production costs, single content and poor interaction of traditional picture books are solved, and low-cost and high-interaction multi-modal picture book production is realized, which stimulates children's learning and imagination and enhances parent-child interaction.

CN120495473APending Publication Date: 2025-08-15湖南工商大学
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510577292.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Traditional picture books are made with labor, with high costs, long cycles, single content, poor matching when fusion of multimodal materials, lack of interaction and immersion experience, and it is difficult to meet the diverse reading needs of modern children.

Method used

AIGC technology is used to generate picture texts, illustrations and audio, combined with AR technology, multimodal interaction logic and personalized customization are designed, and multimodal fusion and dynamic presentation of picture books are realized through GPT-4, DALL·E, StableDiffusion, TTS, ARKit, ARCore and other tools.

Benefits of technology

It reduces the cost of picture book production, enhances visual appeal and interactivity, stimulates children's learning potential and reading interest, expands imagination space, and enhances parent-child interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495473A_ABST
    Figure CN120495473A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence content generation, and discloses an AI GC-based multi-modal picture book material fusion and dynamic presentation method. The method comprises the following specific steps: S1, creating a picture book; s2, illustration matching generation; s3, audio production and addition; s4, preliminarily fusing the materials; s5, interactive logic design; s6, embedding AR elements; and S7, carrying out personalized customization recommendation. According to the method, knowledge picture books are generated by means of AI GC, multi-mode and AR assistance are achieved, abstraction is converted into images, learning potential is stimulated, and knowledge construction, personalized customization, special effect matching, AR and expansion imagination are assisted; a warm bridge is built with the parent-child, the parent-child explores the world together, multiple modes are achieved, AR rich interaction is achieved, the parent-child distance is shortened, the barrier is eliminated, and the parent-child is warm and sweet in light.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence generated content, and specifically provides a multimodal picture book material fusion and dynamic presentation method based on AIGC. Background Art

[0002] In the field of picture book production, the traditional model has long dominated, and its production process is highly dependent on manual work. From the initial story conception and text writing, to the careful drawing of each illustration, and then to the later audio recording and production, all of them require a lot of manpower, material resources and time costs. Professional illustrators charge high fees for drawing illustrations, and the drawing cycle is long. Text creators also need to polish the manuscripts repeatedly, which means that a picture book often takes several months or even longer from the germination of ideas to the release of the finished product. Moreover, due to the limitations of manual creation, the content style of traditional picture books is relatively single, most of them follow established routines, and it is difficult to innovate, which cannot meet the increasingly diverse reading preferences of children.

[0003] With the rapid development of science and technology, AIGC technology has emerged, bringing hope to picture book production. It can rely on powerful algorithms to automatically generate picture book texts, draw exquisite images, and synthesize reading audio, greatly shortening the production cycle and reducing costs. However, the current application of AIGC technology in picture book production is not perfect. When multimodal materials are integrated, the matching degree between text and image, the coordination between audio and picture are poor, and the dynamic presentation process lacks fluency and appeal.

[0004] Furthermore, traditional picture books basically remain at the static reading level, and children can only watch passively, lacking opportunities for interactive communication and immersive experience. This runs counter to the reading demands of modern children who pursue novelty and desire deep participation, and urgently need innovation and change to adapt to the needs of the times. Summary of the Invention

[0005] The purpose of the present invention is to provide a multimodal picture book material fusion and dynamic presentation method based on AIGC to solve the problems raised in the above background technology.

[0006] In order to achieve the above-mentioned purpose, the present invention provides the following technical solution: A multimodal picture book material fusion and dynamic presentation method based on AIGC has the following specific steps:

[0007] S1: Picture book text creation: Use GPT-4 to create interesting and informative texts with good plot and language based on children's cognition and popular topics, to facilitate multimodal development;

[0008] S2: Illustration Matching Generation: Leveraging the image generation models DALL E and StableDiffusion, we accurately transform text descriptions into colorful, stylized illustrations, focusing on visual details and character expressiveness. This ensures that the illustrations closely align with the text, enhancing the visual appeal of the picture book and helping children understand the story.

[0009] S3: Audio Production and Addition: Using TTS technology, select a voice that fits the story atmosphere to read the text, generate clear and smooth audio, and then carefully select soothing or exciting background music, combined with realistic sound effects, to create an immersive feeling, allowing children to immerse themselves in the audio-visual feast of picture books;

[0010] S4: Initial integration of materials: align the generated text, images, and audio precisely according to the story sequence, unify the layout format, adjust the image-text ratio and audio playback nodes, ensure the smooth connection of each modality, form a preliminary complete and coherent picture book prototype, and present the full picture of the story;

[0011] S5: Interaction Logic Design: Plan multimodal interaction paths, set text prompts, image changes, and audio feedback based on user operations, touch characters to trigger dialogues, voice commands to switch scenes, and gestures to scale 3D objects, to ensure smooth content interaction and enhance children's participation and desire to explore;

[0012] S6: AR element embedding: Based on the fusion picture book, AR rendering engines ARKit and ARCore are used to build 3D characters and dynamic scene AR interactive elements, giving virtual content light and shadow effects and physical properties, allowing it to naturally integrate into reality, injecting fresh vitality into the picture book and expanding the imagination space;

[0013] S7: Personalized customized recommendations: Comprehensively consider the user's age, interests, and reading history, select suitable materials to create picture books, set control permissions for parents, filter out bad information, and accurately push personalized picture books to meet diverse needs and help children grow in reading.

[0014] Preferably, the picture book text creation in S1 refers to the use of the natural language generation model GPT-4 to conceive a story framework based on children's cognitive characteristics and popular topics, create picture book texts with rich plots and vivid language, incorporate interesting knowledge, stimulate children's interest in reading, and lay the foundation for subsequent multimodal development.

[0015] Preferably, the specific steps of generating the illustration matching in S2 are as follows:

[0016] Step 1: Text parsing and style positioning: Carefully study the picture book text generated by GPT-4, extract key scenes, characters, and emotional tone information, and determine the illustration style based on the text characteristics and the preferences of the target audience. Targeting the cartoon-like cute style for young children and the fantasy-realistic style for older children, this will provide a clear direction for subsequent image generation;

[0017] Step 2: Accurate Image Generation: The identified style and key text information is fed into the DALL E and Stable Diffusion models. Leveraging the models' powerful image generation capabilities, preliminary illustrations are generated. The prompt word parameters are adjusted multiple times to ensure vibrant colors, reasonable composition, and accurate representation of each scene in the text.

[0018] Step 3: Detail optimization and echo adjustment: Manually review the generated illustrations, focusing on optimizing picture details, strengthening the expressiveness of the characters' expressions and actions to make them vivid and lively, comparing the text plot, checking the echo between the illustrations and the text description, and modifying any mismatches, so that children can quickly associate the corresponding text content just by looking at the illustrations, effectively enhancing the visual appeal of the picture book and helping children understand the story.

[0019] Preferably, the specific steps of adding audio production in S3 are as follows:

[0020] Step 1: Adapting the timbre to the text: Analyze the picture book text in depth to determine the overall style of the story, whether it's a heartwarming fairy tale, an exciting adventure story, or a fantasy-filled myth. Based on these different styles, use the timbre library of TTS technology to carefully select timbre that fits the story's atmosphere. Choose a soft and sweet female voice for a heartwarming story, and a passionate and powerful male voice for an adventure story. This ensures the timbre accurately conveys the story's emotional tone, laying a solid foundation for subsequent audio production.

[0021] Step 2: Generate basic audio: After selecting the appropriate timbre, input the picture book text into the TTS system. By fine-tuning the speech rate, intonation, and pause parameters, the reading rhythm matches the rhythm of the story, generating a clear, smooth, and rhythmic audio file. During this process, special attention should be paid to the handling of polyphonetic characters and unstressed words to avoid reading errors, ensure that children can smoothly understand the text content, and let the story be vividly presented through the sound;

[0022] Step 3: Integrate sound effects and background: Based on the existing reading audio, carefully select corresponding soothing and exciting background music according to different scenes in the story, such as quiet nights, turbulent rivers, and fierce battles, as well as realistic sound effects of wind, rain, and weapon collisions. Use audio editing software to mix the background music, sound effects, and reading audio in a certain proportion. Pay attention to adjusting the volume balance to avoid mutual coverage. Through clever integration, create a sense of immersion, making children feel as if they are in the world depicted in the picture book, and immerse themselves in the audio-visual feast of the picture book.

[0023] Preferably, the initial fusion of materials in S4 refers to taking the text as the basis and combining the corresponding exquisite illustrations. Figure 1They are placed one by one to ensure that the pictures accurately interpret the meaning of the text. At the same time, according to the importance of the plot and the complexity of the pictures on each page, the picture-text ratio is scientifically adjusted to achieve visual balance. In terms of audio, according to the rhythm of the story and the switching of pictures, the playback nodes are accurately set to synchronize the sound and pictures. Only with such careful arrangement can the prototype of the picture book appear and the charm of the story be fully displayed.

[0024] Preferably, the specific steps of the interaction logic design in S5 are as follows:

[0025] Step 1: Interaction Path Planning and Design: A comprehensive analysis of the story content and audience characteristics of the picture book was conducted. Based on different plot segments and interaction needs, a multimodal interaction path encompassing touch, voice, and gesture operations was designed. The text prompts that pop up when touching a character, the resulting dynamic image changes, and the accompanying audio feedback effects were clearly defined. The corresponding relationship between voice commands and scene transitions, as well as the control logic of gestures for 3D objects, were carefully planned.

[0026] Step 2: Interactive effect optimization test: Build an interactive system according to the planned scheme, integrate text, images, and audio materials into it, repeatedly test the actual effects of touch, voice, and gesture operations, observe the children's reactions during the operation, and promptly optimize the text prompts based on feedback to ensure that the language is concise and easy to understand, the image changes are natural and smooth, and the audio feedback is just right, to ensure that the entire interactive process is smooth and unobstructed, and to maximize the children's participation and desire to explore.

[0027] Preferably, the specific steps of embedding the AR element in S6 are as follows:

[0028] Step 1: Build AR interactive elements: Leveraging ARKit and ARCore's professional AR rendering engines, we delved deeply into the key characters and scenes in the picture book. Using 3D modeling technology, we meticulously crafted lifelike 3D characters, striving for perfection in every detail, from their appearance to their clothing. At the same time, we created dynamic scenes that closely aligned with the story's development, ensuring that these elements were designed to both highlight the story's essence and create a unique visual impact.

[0029] Step 2: Optimize the virtual integration effect: After completing the basic construction, start to give the virtual content dynamic lighting effects, simulate the refraction and reflection of light in the real world, and make the 3D characters and dynamic scenes seem to be bathed in real light. Then, according to the laws of physics, add the physical properties of mass, gravity, and collision to it, so that the movement of virtual objects follows the laws of nature. Through repeated debugging and optimization, the virtual content can be seamlessly integrated into the real environment, complementing the physical part of the picture book, opening a door to a fantasy world for children and greatly expanding their imagination.

[0030] Preferably, the specific steps of the personalized recommendation in S7 are as follows:

[0031] Step 1: Preparation for the creation of personalized picture books: Establish a professional data analysis team to collect and organize user age, interests, and reading history information, and use big data algorithms to conduct in-depth mining and analysis of this data. Based on the cognitive characteristics of children of different age groups, accurately identify needs, and based on the analysis results, select suitable elements from a large amount of text, images, and audio materials to lay the foundation for the creation of personalized picture books.

[0032] Step 2: Coordination of picture book push and parental control: After completing the picture book creation, build an intelligent push platform to accurately push the picture book to the target children based on the preliminary analysis data. At the same time, create a dedicated control background for parents, grant them the authority to view and filter the content of the picture book, and use the intelligent filtering system to block violent, vulgar, superstitious and negative information in advance. Parents can adjust the push settings at any time according to the growth of their children to ensure that the pushed picture books can not only meet the children's diverse reading needs and stimulate their interest in reading, but also help children's reading growth under the protection of their parents.

[0033] The beneficial effects of the present invention are as follows:

[0034] 1. This invention uses AIGC technology to accurately generate picture book texts that fit children's cognition, incorporates interesting scientific and historical knowledge, and turns the abstract into the concrete, making learning easy and interesting. Multimodal fusion presentation, such as realistic illustrations and vivid audio, stimulates the senses in all directions, deepens knowledge understanding and memory, and AR technology activates knowledge. Virtual characters interact to answer questions, and dynamic scenes display principles. Children feel like they are in an ocean of knowledge, actively exploring the unknown, greatly stimulating their learning potential, injecting innovative vitality into traditional education, and helping to build a children's knowledge system.

[0035] 2. The fantasy stories and colorful pictures created by AIGC in this invention instantly capture children's attention, and personalized customization can meet diverse preferences. Whether it is a fantasy fairy tale or a science fiction adventure, there is always one that attracts the eyeballs. Combined with audio special effects, the immersive feeling is maximized, and children seem to be stepping into an adventure in another world. With the blessing of AR, virtual playmates accompany them at all times, and the castles and forests in the books are tangible and feelable, allowing children to use their imagination to shuttle through them, expand the boundaries of thinking, and fill their leisure time with infinite creativity. It has become a new favorite of children's entertainment, giving them pure happiness and free imagination.

[0036] 3. The present invention builds a warm bridge through this system. Parents and children can explore the picture book world generated by AIGC together, share the joy of reading, and get closer through interesting knowledge quizzes and story role-playing. The multimodal and AR features make the interaction diverse. Children excitedly tell what they see and feel, and parents guide and inspire them in a timely manner. Emotional communication is natural and smooth. After a busy day, parents and children can sit together and immerse themselves in wonderful stories with the help of the system, creating beautiful memories together, resolving daily barriers, injecting warmth into the family, making parent-child time the warmest companionship on the child's growth path, and strengthening family emotional bonds. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 This is a flow chart of the multimodal picture book material fusion and dynamic presentation method based on AIGC of the present invention. DETAILED DESCRIPTION

[0038] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0039] like Figure 1 As shown, the embodiment of the present invention provides a multimodal picture book material fusion and dynamic presentation method based on AIGC, and the specific steps are as follows:

[0040] S1: Picture book text creation: Use GPT-4 to create interesting and informative texts with good plot and language based on children's cognition and popular topics, to facilitate multimodal development;

[0041] S2: Illustration Matching Generation: Using the image generation models DALL E and StableDisfusion, we accurately transform text descriptions into colorful and stylish illustrations, focusing on image details and character expressiveness. This ensures that the illustrations closely align with the text, enhancing the visual appeal of the picture book and helping children understand the story.

[0042] S3: Audio Production and Addition: Using TTS technology, select a voice that fits the story atmosphere to read the text, generate clear and smooth audio, and then carefully select soothing or exciting background music, combined with realistic sound effects, to create an immersive feeling, allowing children to immerse themselves in the audio-visual feast of picture books;

[0043] S4: Initial integration of materials: align the generated text, images, and audio precisely according to the story sequence, unify the layout format, adjust the image-text ratio and audio playback nodes, ensure the smooth connection of each modality, form a preliminary complete and coherent picture book prototype, and present the full picture of the story;

[0044] S5: Interaction Logic Design: Plan multimodal interaction paths, set text prompts, image changes, and audio feedback based on user operations, touch characters to trigger dialogues, voice commands to switch scenes, and gestures to scale 3D objects, to ensure smooth content interaction and enhance children's participation and desire to explore;

[0045] S6: AR element embedding: Based on the fusion picture book, AR rendering engines ARKit and ARCore are used to build 3D characters and dynamic scene AR interactive elements, giving virtual content light and shadow effects and physical properties, allowing it to naturally integrate into reality, injecting fresh vitality into the picture book and expanding the imagination space;

[0046] S7: Personalized customized recommendations: Comprehensively consider the user's age, interests, and reading history, select suitable materials to create picture books, set control permissions for parents, filter out bad information, and accurately push personalized picture books to meet diverse needs and help children grow in reading.

[0047] Among them, the picture book text creation in S1 refers to the use of the natural language generation model GPT-4 to conceive the story framework based on children's cognitive characteristics and popular topics, create picture book texts with rich plots and vivid language, incorporate interesting knowledge, stimulate children's interest in reading, and lay the foundation for subsequent multimodal development.

[0048] The core lies in the clever use of the powerful tool GPT-4, in-depth analysis of children's cognitive laws, grasping popular trendy themes, carefully outlining the story, weaving the text with interesting plots and lively language, incorporating knowledge Easter eggs, arousing children's enthusiasm for reading, and embarking on a multimodal journey.

[0049] The specific steps for generating the illustration matching in S2 are as follows:

[0050] Step 1: Text parsing and style positioning: Carefully study the picture book text generated by GPT-4, extract key scenes, characters, and emotional tone information, and determine the illustration style based on the text characteristics and the preferences of the target audience. Targeting the cartoon-like cute style for young children and the fantasy-realistic style for older children, this will provide a clear direction for subsequent image generation;

[0051] Step 2: Accurate Image Generation: The identified style and key text information is fed into the DALL E and Stable Diffusion models. Leveraging the models' powerful image generation capabilities, preliminary illustrations are generated. The prompt word parameters are adjusted multiple times to ensure vibrant colors, reasonable composition, and accurate representation of each scene in the text.

[0052] Step 3: Detail optimization and echo adjustment: Manually review the generated illustrations, focusing on optimizing picture details, strengthening the expressiveness of the characters' expressions and actions to make them vivid and lively, comparing the text plot, checking the echo between the illustrations and the text description, and modifying any mismatches, so that children can quickly associate the corresponding text content just by looking at the illustrations, effectively enhancing the visual appeal of the picture book and helping children understand the story.

[0053] The first step is text analysis and style positioning. The GPT-4 text is studied, key elements are extracted, and the illustration style is locked according to the audience's preferences. The cute style and fantasy realistic style are adapted for toddlers and older children respectively to guide the direction of image generation; then accurate image generation is carried out, and the style and key information are input into the model to adjust the colors to be bright and the composition to be precise; finally, the details are optimized and the echoes are adjusted, and manual review is carried out to optimize the details, verify the echoes, enhance the visual appeal and aid understanding.

[0054] The specific steps for adding audio production in S3 are as follows:

[0055] Step 1: Adapting the timbre to the text: Analyze the picture book text in depth to determine the overall style of the story, whether it's a heartwarming fairy tale, an exciting adventure story, or a fantasy-filled myth. Based on these different styles, use the timbre library of TTS technology to carefully select timbre that fits the story's atmosphere. Choose a soft and sweet female voice for a heartwarming story, and a passionate and powerful male voice for an adventure story. This ensures the timbre accurately conveys the story's emotional tone, laying a solid foundation for subsequent audio production.

[0056] Step 2: Generate basic audio: After selecting the appropriate timbre, input the picture book text into the TTS system. By fine-tuning the speech rate, intonation, and pause parameters, the reading rhythm matches the rhythm of the story, generating a clear, smooth, and rhythmic audio file. During this process, special attention should be paid to the handling of polyphonetic characters and unstressed words to avoid reading errors, ensure that children can smoothly understand the text content, and let the story be vividly presented through the sound;

[0057] Step 3: Integrate sound effects and background: Based on the existing reading audio, carefully select corresponding soothing and exciting background music according to different scenes in the story, such as quiet nights, turbulent rivers, and fierce battles, as well as realistic sound effects of wind, rain, and weapon collisions. Use audio editing software to mix the background music, sound effects, and reading audio in a certain proportion. Pay attention to adjusting the volume balance to avoid mutual coverage. Through clever integration, create a sense of immersion, making children feel as if they are in the world depicted in the picture book, and immerse themselves in the audio-visual feast of the picture book.

[0058] First, the timbre adapts to the text, deeply analyzes the picture book text, and accurately determines the story style, such as a warm fairy tale, an exciting adventure, or a fantasy myth; accordingly, it shuttles through the TTS sound library to find a soft female voice for a warm atmosphere and a passionate male voice for an adventurous journey, laying the emotional tone; second, it generates basic audio, selects the timbre and inputs the text into the TTS system, finely controls the speed, tone, and pauses, and rigorously handles details such as polyphones to make the reading smooth and rhythmic, and clearly convey the story; third, it integrates the sound background, selects background music and realistic sound effects according to the story scene, uses software to mix them in proportion, balances the volume, creates an immersive feeling, and invites children to immerse themselves in the audio-visual world of picture books.

[0059] The initial fusion of materials in S4 refers to taking text as the basis and combining the corresponding exquisite illustrations. Figure 1 They are placed one by one to ensure that the pictures accurately interpret the meaning of the text. At the same time, according to the importance of the plot and the complexity of the pictures on each page, the picture-text ratio is scientifically adjusted to achieve visual balance. In terms of audio, according to the rhythm of the story and the switching of pictures, the playback nodes are accurately set to synchronize the sound and pictures. Only with such careful arrangement can the prototype of the picture book appear and the charm of the story be fully displayed.

[0060] The exquisite illustrations are precisely matched with the text, allowing the pictures to accurately convey the meaning of the text. According to the plot and picture conditions of each page, the proportion of pictures and text is scientifically adjusted to create a visual balance. In terms of audio, the playback nodes are accurately set according to the rhythm of the story and the transition of the pictures to achieve synchronization of sound, picture and text. After this careful arrangement, the prototype of the picture book is born, fully demonstrating the charm of the story.

[0061] The specific steps of the interaction logic design in S5 are as follows:

[0062] Step 1: Interaction Path Planning and Design: A comprehensive analysis of the story content and audience characteristics of the picture book was conducted. Based on different plot segments and interaction needs, a multimodal interaction path encompassing touch, voice, and gesture operations was designed. The text prompts that pop up when touching a character, the resulting dynamic image changes, and the accompanying audio feedback effects were clearly defined. The corresponding relationship between voice commands and scene transitions, as well as the control logic of gestures for 3D objects, were carefully planned.

[0063] Step 2: Interactive effect optimization test: Build an interactive system according to the planned scheme, integrate text, images, and audio materials into it, repeatedly test the actual effects of touch, voice, and gesture operations, observe the children's reactions during the operation, and promptly optimize the text prompts based on feedback to ensure that the language is concise and easy to understand, the image changes are natural and smooth, and the audio feedback is just right, to ensure that the entire interactive process is smooth and unobstructed, and to maximize the children's participation and desire to explore.

[0064] To plan and design the interactive path, we conducted in-depth research on the picture book stories and audience characteristics, closely followed the plot and interactive demands, and carefully created an interactive path that covers multiple operation methods. We refined various types of feedback triggered by touch, voice, and gestures, such as the linkage between pictures, texts, and sounds after touching the characters. Then we conducted interactive effect optimization tests. We built the system according to the plan, integrated the materials, repeatedly practiced, and optimized the details according to the children's reactions to ensure a smooth process and stimulate children's enthusiasm for exploration.

[0065] The specific steps of embedding the AR element in S6 are as follows:

[0066] Step 1: Build AR interactive elements: Leveraging ARKit and ARCore's professional AR rendering engines, we delved deeply into the key characters and scenes in the picture book. Using 3D modeling technology, we meticulously crafted lifelike 3D characters, striving for perfection in every detail, from their appearance to their clothing. At the same time, we created dynamic scenes that closely aligned with the story's development, ensuring that these elements were designed to both highlight the story's essence and create a unique visual impact.

[0067] Step 2: Optimize the virtual integration effect: After completing the basic construction, start to give the virtual content dynamic lighting effects, simulate the refraction and reflection of light in the real world, and make the 3D characters and dynamic scenes seem to be bathed in real light. Then, according to the laws of physics, add the physical properties of mass, gravity, and collision to it, so that the movement of virtual objects follows the laws of nature. Through repeated debugging and optimization, the virtual content can be seamlessly integrated into the real environment, complementing the physical part of the picture book, opening a door to a fantasy world for children and greatly expanding their imagination.

[0068] The first step is to build AR interactive elements. Relying on ARKit and ARCore engines, we deeply explore the key elements of picture books and use 3D modeling to carefully sculpt 3D characters. The appearance and clothing are meticulously designed, and the dynamic scenes that fit the story are matched to show the charm of the story and attract attention. In order to optimize the virtual integration effect, realistic light and shadow are added to the virtual content, simulating the refraction and reflection of light, and giving mass, gravity and other characteristics according to the laws of physics. After repeated debugging, it is perfectly integrated with reality and picture books, opening the door to fantasy and letting children's imagination fly.

[0069] The specific steps of the personalized customization recommendation in S7 are as follows:

[0070] Step 1: Preparation for the creation of personalized picture books: Establish a professional data analysis team to collect and organize user age, interests, and reading history information, and use big data algorithms to conduct in-depth mining and analysis of this data. Based on the cognitive characteristics of children of different age groups, accurately identify needs, and based on the analysis results, select suitable elements from a large amount of text, images, and audio materials to lay the foundation for the creation of personalized picture books.

[0071] Step 2: Coordination of picture book push and parental control: After completing the picture book creation, build an intelligent push platform to accurately push the picture book to the target children based on the preliminary analysis data. At the same time, create a dedicated control background for parents, grant them the authority to view and filter the content of the picture book, and use the intelligent filtering system to block violent, vulgar, superstitious and negative information in advance. Parents can adjust the push settings at any time according to the growth of their children to ensure that the pushed picture books can not only meet the children's diverse reading needs and stimulate their interest in reading, but also help children's reading growth under the protection of their parents.

[0072] First, in the preparation for the creation of personalized picture books, a professional team collects information such as age, interests, and reading history, analyzes it with big data algorithms, and accurately locates children based on their cognitive differences; then, they filter out suitable elements from massive materials to lay the foundation for customized picture books. Second, picture book push is coordinated with parental control. After the creation is completed, an intelligent platform is built for precise push, and a control console and filtering system are set up for parents to filter content, block bad information, and adjust settings as needed to protect children's reading.

[0073] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0074] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A multimodal picture book material fusion and dynamic presentation method based on AIGC, characterized by: The specific steps of the AIGC-based multimodal picture book material fusion and dynamic presentation method are as follows: S1: Picture book text creation: Use GPT-4 to create interesting and informative texts with good plot and language based on children's cognition and popular topics, to facilitate multimodal development; S2: Illustration Matching Generation: Leveraging the image generation models DALL E and StableDiffusion, we accurately transform text descriptions into colorful, stylized illustrations, focusing on visual details and character expressiveness. This ensures that the illustrations closely align with the text, enhancing the visual appeal of the picture book and helping children understand the story. S3: Audio Production and Addition: Using TTS technology, select a voice that fits the story atmosphere to read the text, generate clear and smooth audio, and then carefully select soothing or exciting background music, combined with realistic sound effects, to create an immersive feeling, allowing children to immerse themselves in the audio-visual feast of picture books; S4: Initial integration of materials: align the generated text, images, and audio precisely according to the story sequence, unify the layout format, adjust the image-text ratio and audio playback nodes, ensure the smooth connection of each modality, form a preliminary complete and coherent picture book prototype, and present the full picture of the story; S5: Interaction Logic Design: Plan multimodal interaction paths, set text prompts, image changes, and audio feedback based on user operations, touch characters to trigger dialogues, voice commands to switch scenes, and gestures to scale 3D objects, to ensure smooth content interaction and enhance children's participation and desire to explore; S6: AR element embedding: Based on the fusion picture book, AR rendering engines ARKit and ARCore are used to build 3D characters and dynamic scene AR interactive elements, giving virtual content light and shadow effects and physical properties, allowing it to naturally integrate into reality, injecting fresh vitality into the picture book and expanding the imagination space; S7: Personalized customized recommendations: Comprehensively consider the user's age, interests, and reading history, select suitable materials to create picture books, set control permissions for parents, filter out bad information, and accurately push personalized picture books to meet diverse needs and help children grow in reading.

2. The method for fusion and dynamic presentation of multimodal picture book materials based on AIGC according to claim 1, characterized in that: The picture book text creation in S1 refers to the use of the natural language generation model GPT-4 to conceive a story framework based on children's cognitive characteristics and popular topics, create picture book texts with rich plots and vivid language, incorporate interesting knowledge, stimulate children's interest in reading, and lay the foundation for subsequent multimodal development.

3. The method for fusion and dynamic presentation of multimodal picture book materials based on AIGC according to claim 1, characterized in that: The specific steps for generating the illustration matching in S2 are as follows: Step 1: Text parsing and style positioning: Carefully study the picture book text generated by GPT-4, extract key scenes, characters, and emotional tone information, and determine the illustration style based on the text characteristics and the preferences of the target audience. Targeting the cartoon-like cute style for young children and the fantasy-realistic style for older children, this will provide a clear direction for subsequent image generation; Step 2: Accurate Image Generation: The identified style and key text information is fed into the DALL E and Stable Diffusion models. Leveraging the models' powerful image generation capabilities, preliminary illustrations are generated. The prompt word parameters are adjusted multiple times to ensure vibrant colors, reasonable composition, and accurate representation of each scene in the text. Step 3: Detail optimization and echo adjustment: Manually review the generated illustrations, focusing on optimizing picture details, strengthening the expressiveness of the characters' expressions and actions to make them vivid and lively, comparing the text plot, checking the echo between the illustrations and the text description, and modifying any mismatches, so that children can quickly associate the corresponding text content just by looking at the illustrations, effectively enhancing the visual appeal of the picture book and helping children understand the story.

4. The method for fusion and dynamic presentation of multimodal picture book materials based on AIGC according to claim 1, characterized in that: The specific steps for adding audio production in S3 are as follows: Step 1: Adapting the timbre to the text: Analyze the picture book text in depth to determine the overall style of the story, whether it's a heartwarming fairy tale, an exciting adventure story, or a fantasy-filled myth. Based on these different styles, use the timbre library of TTS technology to carefully select timbre that fits the story's atmosphere. Choose a soft and sweet female voice for a heartwarming story, and a passionate and powerful male voice for an adventure story. This ensures the timbre accurately conveys the story's emotional tone, laying a solid foundation for subsequent audio production. Step 2: Generate basic audio: After selecting the appropriate timbre, input the picture book text into the TTS system. By fine-tuning the speech rate, intonation, and pause parameters, the reading rhythm matches the rhythm of the story, generating a clear, smooth, and rhythmic audio file. During this process, special attention should be paid to the handling of polyphonetic characters and unstressed words to avoid reading errors, ensure that children can smoothly understand the text content, and let the story be vividly presented through the sound; Step 3: Integrate sound effects and background: Based on the existing reading audio, carefully select corresponding soothing and exciting background music according to different scenes in the story, such as quiet nights, turbulent rivers, and fierce battles, as well as realistic sound effects of wind, rain, and weapon collisions. Use audio editing software to mix the background music, sound effects, and reading audio in a certain proportion. Pay attention to adjusting the volume balance to avoid mutual coverage. Through clever integration, create a sense of immersion, making children feel as if they are in the world depicted in the picture book, and immerse themselves in the audio-visual feast of the picture book.

5. The method for fusion and dynamic presentation of multimodal picture book materials based on AIGC according to claim 1, characterized in that: The initial integration of materials in S4 refers to placing the corresponding exquisite illustrations one by one based on the text to ensure that the pictures accurately interpret the meaning of the text. At the same time, according to the importance of the plot and the complexity of the pictures on each page, the proportion of pictures and text is scientifically adjusted to achieve visual balance. In terms of audio, according to the rhythm of the story and the switching of pictures, the playback nodes are accurately set to synchronize the sound and pictures. Only with such careful arrangement can the prototype of the picture book appear and the charm of the story be fully displayed.

6. The method for fusion and dynamic presentation of multimodal picture book materials based on AIGC according to claim 1, characterized in that: The specific steps of the interaction logic design in S5 are as follows: Step 1: Interaction Path Planning and Design: A comprehensive analysis of the story content and audience characteristics of the picture book was conducted. Based on different plot segments and interaction needs, a multimodal interaction path encompassing touch, voice, and gesture operations was designed. The text prompts that pop up when touching a character, the resulting dynamic image changes, and the accompanying audio feedback effects were clearly defined. The corresponding relationship between voice commands and scene transitions, as well as the control logic of gestures for 3D objects, were carefully planned. Step 2: Interactive effect optimization test: Build an interactive system according to the planned scheme, integrate text, images, and audio materials into it, repeatedly test the actual effects of touch, voice, and gesture operations, observe the children's reactions during the operation, and promptly optimize the text prompts based on feedback to ensure that the language is concise and easy to understand, the image changes are natural and smooth, and the audio feedback is just right, to ensure that the entire interactive process is smooth and unobstructed, and to maximize the children's participation and desire to explore.

7. The method for fusion and dynamic presentation of multimodal picture book materials based on AIGC according to claim 1, characterized in that: The specific steps for embedding AR elements in S6 are as follows: Step 1: Build AR interactive elements: Leveraging ARKit and ARCore's professional AR rendering engines, we delved deeply into the key characters and scenes in the picture book. Using 3D modeling technology, we meticulously crafted lifelike 3D characters, striving for perfection in every detail, from their appearance to their clothing. At the same time, we created dynamic scenes that closely aligned with the story's development, ensuring that these elements were designed to both highlight the story's essence and create a unique visual impact. Step 2: Optimize the virtual integration effect: After completing the basic construction, start to give the virtual content dynamic lighting effects, simulate the refraction and reflection of light in the real world, and make the 3D characters and dynamic scenes seem to be bathed in real light. Then, according to the laws of physics, add the physical properties of mass, gravity, and collision to it, so that the movement of virtual objects follows the laws of nature. Through repeated debugging and optimization, the virtual content can be seamlessly integrated into the real environment, complementing the physical part of the picture book, opening a door to a fantasy world for children and greatly expanding their imagination.

8. The method for fusion and dynamic presentation of multimodal picture book materials based on AIGC according to claim 1, characterized in that: The specific steps for personalized customization recommendation in S7 are as follows: Step 1: Preparation for the creation of personalized picture books: Establish a professional data analysis team to collect and organize user age, interests, and reading history information, and use big data algorithms to conduct in-depth mining and analysis of this data. Based on the cognitive characteristics of children of different age groups, accurately identify needs, and based on the analysis results, select suitable elements from a large amount of text, images, and audio materials to lay the foundation for the creation of personalized picture books. Step 2: Coordination of picture book push and parental control: After completing the picture book creation, build an intelligent push platform to accurately push the picture book to the target children based on the preliminary analysis data. At the same time, create a dedicated control background for parents, grant them the authority to view and filter the content of the picture book, and use the intelligent filtering system to block violent, vulgar, superstitious and negative information in advance. Parents can adjust the push settings at any time according to the growth of their children to ensure that the pushed picture books can not only meet the children's diverse reading needs and stimulate their interest in reading, but also help children's reading growth under the protection of their parents.

Citation Information

Cited By

  • Intelligent interaction and self-adaptive prompting method and system for cognitive assessment of old people

    CN122117417A