Picture book interaction method and picture book content artificial intelligence generation method and system
By combining a picture book interactive system with an AI engine, the operation and data security issues of electronic picture books on smart TVs have been solved, enabling the generation of personalized picture book content and secure interaction, thus enriching the picture book reading experience on TV.
Patent Information
- Application Number
- CN202510926892.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-06-30
- Filing Date
- 2025-07-07
- Publication Date
- 2025-11-25
AI Technical Summary
Existing smart TV voice remote controls and their voice assistants cannot directly operate electronic picture books, and there are technical compatibility and information sharing security risks. This results in electronic picture book software needing extensive adaptation for different TV manufacturers, and business data is also at risk of security vulnerabilities.
The system employs a picture book interactive system, including picture book application software and server-side system. It utilizes the front-end SDK and back-end service system of the picture book interactive AI engine, combined with an interactive picture book media library and a picture book content artificial intelligence generation system. It generates personalized and customized picture book content through voice interaction, and realizes picture book interaction and content generation independently of the TV's intelligent voice assistant.
It enables a rich e-book reading experience and immersive interaction on smart TVs, reduces the adaptation work with TV manufacturers, and protects the business data security of picture book application software.
Smart Images

Figure CN121008684A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of artificial intelligence, and relates to a picture book interaction method and a picture book content artificial intelligence generation method and system. BACKGROUND
[0002] Electronic picture books have become the first choice for parents and children due to their portability, strong interactivity, and upgradability. For example, H5 technology is used to develop interactive electronic picture books, which enhance the reading experience through clicking, dragging, and other operations. Currently, electronic picture books are generally provided to child readers on tablets. However, tablets have smaller screens and are closer to the eyes, causing much greater stimulation to the ciliary muscle than television. Studies have shown that watching cartoons on a tablet is equivalent to 10 times the stimulation to the ciliary muscle as watching television. Although current televisions have the hardware and software conditions to install and run electronic picture books, the operation of the television remote control is not as convenient and flexible as that of a tablet, limiting the use of televisions as a reading carrier for electronic picture books.
[0003] With the popularization and application of artificial intelligence and large model technology, current televisions of multiple brands have realized television operation based on large models through voice remote controls. These operations are usually implemented through voice assistants running on the television. The voice assistant integrates large model technology to recognize user intent, enabling the voice remote control to perform all the functions of a regular remote control. After adapting to installed television APP application software, the voice remote control also supports starting and logging into the APP application software (refer to patent CN119402690A). However, existing televisions with voice remote controls and voice assistants based on large model technology cannot directly operate electronic picture books running on the television, and there are also issues such as technical adaptation and information sharing.
[0004] Based on the existing voice remote control of smart televisions and the voice assistant supporting large model technology, electronic picture book interaction is realized, and personalized and customized picture book functions are provided. However, there are still the following problems:
[0005] (1) The voice remote control of a smart television and the voice assistant supporting large model technology are functions provided by the television system itself. They are mainly used to interact with the functions of the television's own or pre-installed software. However, electronic picture books are not usually part of the television's own or pre-installed software. Therefore, different television manufacturers need to be adapted to realize the interaction of personalized and customized functions of electronic picture books. This obviously requires electronic picture book software and its service providers to perform a large amount of adaptation work for different television manufacturers.
[0006] (2) Electronic picture book software and its service providers are not part of the television's own or pre-installed software. The unique business data of this software and its service providers poses a security risk of information sharing. SUMMARY
[0007] In order to overcome at least one of the deficiencies of the prior art, the present application provides a picture book interaction method and a picture book content artificial intelligence generation method and system.
[0008] In order to achieve the above-mentioned purpose, the present application adopts the following technical solution: a picture book interaction system, comprising a picture book application software and a server system, the picture book application software comprising a picture book interaction AI engine front-end SDK and a role-playing mode program, the server system comprising a picture book interaction AI engine back-end service system, an interactive picture book media asset library system and a picture book content artificial intelligence generation system,
[0009] The picture book interaction AI engine front-end SDK is used to realize picture book interaction, picture book media asset data calling and picture book content artificial intelligence generation system to generate a picture book;
[0010] The role-playing mode program: in the reading mode, according to the voice request of the user, the picture book interaction AI engine front-end SDK is called internally and initiates a picture book request to the picture book content artificial intelligence generation system;
[0011] The picture book interaction AI engine back-end service system: provides cloud AI service capability for the picture book interaction AI engine front-end SDK, provides artificial intelligence analysis and processing for user interaction request, and feeds back the analysis and processing result to the picture book interaction AI engine front-end SDK;
[0012] The interactive picture book media asset library system: is an interactive electronic picture book media resource metadata system, containing picture book image data and audio data, and provides picture book data calling service for the picture book interaction AI engine front-end SDK;
[0013] The picture book content artificial intelligence generation system: generates picture book metadata of a new picture book based on the user request role by artificial intelligence, and pushes it to the picture book application software, the picture book application software receives the picture book metadata and displays and lays out to show.
[0014] Further, the picture book content artificial intelligence generation system includes a large model pre-trained through a fairy tale corpus, a text-to-picture model fine-tuned through a children's picture book dataset, a text-to-audio model fine-tuned through a children's picture book dataset, and an interactive picture book arrangement system. The large model pre-trained through the fairy tale corpus generates a new picture book story text based on a user requested role. The text-to-picture model fine-tuned through the children's picture book dataset inputs the new picture book story text generated by the large model and outputs a new picture book image. The text-to-audio model fine-tuned through the children's picture book dataset inputs the new picture book story text generated by the large model and outputs a new picture book audio. The interactive picture book arrangement system combines the picture book image and the picture book audio to generate a new picture book story, automatically assembles the new picture book story into new picture book metadata, and pushes the picture book metadata to a picture book application software.
[0015] Further, the picture book interactive AI engine front-end SDK and the voice terminal realize voice interaction. The voice terminal receives a user's voice request and analyzes and processes the voice request. The processed data is transmitted to the picture book application software.
[0016] Further, the voice terminal includes a voice terminal, a carrier, and a speech recognition SDK. The voice terminal receives a user's voice request and transmits the voice request to the carrier. After noise reduction processing by the carrier, the voice request is transmitted to the speech recognition SDK. The speech recognition SDK extracts an interactive request voice TTS text and a voice sample in the voice request and transmits the interactive request voice TTS text and the voice sample to the picture book interactive AI engine front-end SDK.
[0017] Further, the picture book interactive AI engine front-end SDK includes a picture book interactive AI capability API interface. The picture book interactive AI engine front-end SDK communicates and cooperates with the picture book interactive AI engine back-end service system through the picture book interactive AI capability API interface, analyzes and processes a user's interactive request using artificial intelligence, and feeds back the analysis and processing results to the picture book interactive AI engine front-end SDK.
[0018] A picture book interaction method includes
[0019] Receiving a picture book reading instruction input by a user to enter a reading mode.
[0020] In the reading mode, receiving a voice request and recognizing the voice request, and opening a specific picture book or entering a role-playing mode based on a picture book reading interaction method.
[0021] In the role-playing mode, generating picture book metadata of a new picture book based on a user requested role through a picture book content artificial intelligence generation method, reading the picture book metadata, and maintaining voice interaction with the user to make the user enter a role-playing picture book.
[0022] Further, the picture book reading interaction method includes the following steps:
[0023] 1) The picture book application software receives the picture book reading instructions input by the user, and the picture book interactive AI engine front-end SDK submits the TTS text and the voice sample to the picture book interactive AI engine back-end service system through the picture book interactive AI capability API interface;
[0024] 2) The picture book interactive AI engine back-end service system analyzes and processes the TTS text and the voice sample, and returns the processed data to the picture book interactive AI engine front-end SDK;
[0025] 3) The picture book interactive AI engine front-end SDK initiates a picture book data calling request to the interactive picture book media asset library system according to the received data;
[0026] 4) The picture book media asset library system submits the picture book metadata to the interactive picture book arrangement system according to the picture book data calling request;
[0027] 5) The interactive picture book arrangement system returns the picture book metadata to the picture book interactive AI engine front-end SDK in the picture book application software;
[0028] 6) The picture book interactive AI engine front-end SDK performs page display based on the picture book application software according to specific data, and submits the reply text based on the picture book interactive AI engine to the speech recognition SDK for conversion into a voice response to the user.
[0029] Further, the method of receiving the picture book reading instructions input by the user: the voice terminal receives the voice of the user and transmits it to the carrier, and transmits it to the speech recognition SDK after noise reduction processing by the carrier, the speech recognition SDK extracts the interactive request voice TTS text and the voice sample in the voice request, and transmits the interactive request voice TTS text and the voice sample to the picture book application software.
[0030] A picture book content artificial intelligence generation method, comprising the following steps:
[0031] 1) The user selects a picture book role through voice under the voice guidance prompt of the role-playing mode program;
[0032] 2) The picture book application software acquires the interactive request voice TTS text and the voice sample through the speech recognition SDK and submits the TTS text and the voice sample to the picture book interactive AI engine front-end SDK;
[0033] 3) The picture book interactive AI engine front-end SDK submits the TTS text and the voice sample to the picture book interactive AI engine back-end service system through the picture book interactive AI capability API interface;
[0034] 4) The picture book interaction AI engine backend service system analyzes and processes the TTS text and voice samples, and returns the processed data to the picture book interaction AI engine frontend SDK;
[0035] 5) The picture book interaction AI engine frontend SDK forwards the received voiceprint recognition, picture book interaction response processing data, and user role selection information to the role-playing mode program, which forwards the received data to the large model pre-trained by the fairy tale corpus;
[0036] 6) The large model pre-trained by the fairy tale corpus analyzes and processes the received data to generate three pieces of text data: ① new picture book story text generated by the large model, ② image element text of the text generated by the large model, and ③ audio element text of the text generated by the large model;
[0037] 7) The new picture book story text generated by the large model pre-trained by the fairy tale corpus is output to the interactive picture book arrangement system;
[0038] 8) The image element text of the text generated by the large model pre-trained by the fairy tale corpus is output to the text-to-image model fine-tuned by the children's picture book dataset;
[0039] 9) The audio element text of the text generated by the large model pre-trained by the fairy tale corpus is output to the text-to-audio model fine-tuned by the children's picture book dataset;
[0040] 10) The text-to-image model fine-tuned by the children's picture book dataset generates picture book split image according to the generated image element text of the text;
[0041] 11) The text-to-audio model fine-tuned by the children's picture book dataset generates picture book split audio according to the generated audio element text of the text;
[0042] 12) The interactive picture book arrangement system automatically arranges the picture book story split image V and the picture book story split audio W according to the received generated content according to the picture book story script number S, the picture book story scene number T, and the picture book story split number U, to generate picture book metadata that can be read by picture book application software;
[0043] 13) The picture book application software reads the automatically arranged picture book metadata, and at the same time responds to the user's voice interaction request through the picture book interaction AI engine frontend SDK and maintains interaction with the user through the voice recognition SDK in audio.
[0044] Further, the generated new picture book story text contains the following key information classification generation:
[0045] Picture book story outline: M + picture book story script: N + picture book story scene: O + picture book character: P + picture book scene and character action: Q + character dialogue: R;
[0046] Picture book story outline M takes the summary text of picture book story script N according to the development order of the story;
[0047] Picture book story script N takes the description text of picture book story scene O according to the development order of the story;
[0048] Picture book story scene O takes the picture book story scene description text;
[0049] Picture book character P takes the character text contained in story scene O;
[0050] Picture book scene and character action Q takes the scene and character action text contained in story scene O;
[0051] Character dialogue R takes the character dialogue text contained in story scene O.
[0052] In summary, the beneficial effects of the present application are:
[0053] 1) The present application is based on the existing intelligent television voice recognition function, and the method and system of the present application realize interactive rich electronic picture book reading and immersive interactive picture book experience based on artificial intelligence generation.
[0054] 2) The present application changes the voice recognition artificial intelligence processing from the current television voice assistant to the picture book interactive AI engine front-end SDK and the picture book interactive AI engine back-end service system of the present application, thereby being independent of the intelligent voice assistant artificial intelligence processing capability of the existing television (set-top box), reducing the intelligent voice assistant adaptation work with the television manufacturer, and realizing business data security protection of the picture book application software. BRIEF DESCRIPTION OF DRAWINGS
[0055] Figure 1 is the overall flowchart of the present application.
[0056] Figure 2 is the picture book interactive system schematic diagram of the present application.
[0057] Figure 3 is the picture book reading interactive method flowchart of the reading mode of the present application.
[0058] Figure 4 is the picture book content artificial intelligence generation method flowchart of the present application. DETAILED DESCRIPTION
[0059] Following make the specific concrete example explain the embodiment of the present application, the person skilled in the art can be easily understood from the present application disclosed in the content of the present application Other advantages and effects.The present application can also be implemented or applied by another different specific implementation, the details in the specification can also be based on different views and applications, without departing from the spirit of the present application, various modifications or changes are made.The need to explain that the following examples and the features in the examples can be combined with each other without conflict.
[0060] Need to explain that the following examples provided in the figure only illustrates the basic concept of the present application, and the figure only shows the components related to the present application, not according to the actual implementation of the number of components, shape and size drawing, the actual implementation of each component type, quantity and proportion can be a kind of arbitrary change, and its component layout type may be more complex.
[0061] All directional indications in the embodiment of the present application (such as up, down, left, right, front, back, transverse, longitudinal...) are only used to explain the relative position relationship, motion condition and so on between the components in a certain posture, if the specific posture changes, the directional indication also changes accordingly.
[0062] Due to installation error and other reasons, the parallel relationship referred to in the embodiment of the present application may be actually an approximate parallel relationship, and the vertical relationship may be actually an approximate vertical relationship.
[0063] Embodiment one:
[0064] As shown in Figures 1-2 A picture book interactive system, including picture book application software and server system, picture book application software includes picture book interactive AI engine front-end SDK and role playing mode program, server system includes picture book interactive AI engine back-end service system, interactive picture book media asset library system and picture book content artificial intelligence generation system;
[0065] Picture book interactive AI engine front-end SDK: for realizing picture book interaction, picture book media data calling and picture book content artificial intelligence generation system generating picture book.
[0066] Role playing mode program: in reading mode, according to the voice request of user, picture book interactive AI engine front-end SDK is called internally and initiates picture book request to picture book content artificial intelligence generation system.
[0067] Picture book interactive AI engine back-end service system: for picture book interactive AI engine front-end SDK to provide cloud AI service capability, provide artificial intelligence analysis and processing to user interactive request, and feedback analysis and processing result to picture book interactive AI engine front-end SDK.
[0068] Interactive picture book media asset library system: an interactive electronic picture book media resource metadata system containing picture book image data and audio data, providing picture book data calling services for the front-end SDK of the picture book interactive AI engine.
[0069] Picture book content AI generation system: an AI system that generates picture book metadata for new picture books based on user requests for characters and pushes them to picture book application software, which receives and displays the picture book metadata.
[0070] The picture book content AI generation system includes a large model pre-trained on a fairy tale corpus, a text-to-image model fine-tuned on a children's picture book dataset, a text-to-audio model fine-tuned on a children's picture book dataset, and an interactive picture book arrangement system. The large model pre-trained on the fairy tale corpus generates new picture book story texts based on user requests for characters. The text-to-image model fine-tuned on the children's picture book dataset inputs and outputs new picture book images based on the new picture book story texts generated by the large model. The text-to-audio model fine-tuned on the children's picture book dataset inputs and outputs new picture book audios based on the new picture book story texts generated by the large model. The interactive picture book arrangement system combines the picture book images and audios to generate new picture book stories, automatically compiles them into new picture book metadata, and pushes the metadata to the picture book application software.
[0071] Large model pre-trained on fairy tale corpus: a large language model pre-trained on a fairy tale corpus, which generates new picture book story texts based on user requests for characters in combination with the requests of the picture book interactive AI engine front-end SDK and the data features of the user's current reading of picture books.
[0072] Interactive picture book arrangement system: supports recommending and retrieving non-AI-generated ordinary interactive electronic picture book data based on user characteristics and requests, supports automatically compiling new picture book metadata from new picture book stories generated by combining AI-generated picture book images and audios, and pushes the recommended, retrieved, and synthesized picture book metadata to the picture book application software.
[0073] The integration of the system is as follows:
[0074] The picture book interactive AI engine front-end SDK and the voice terminal realize voice interaction. The voice terminal receives and analyzes user voice requests and transmits the processed data to the picture book application software.
[0075] The voice terminal includes a voice terminal, a carrier, and a voice recognition SDK. The voice terminal receives a user's voice request through a conventional method, and transmits the voice request to the carrier. After noise reduction processing by the carrier, the voice request is transmitted to the voice recognition SDK. The voice recognition SDK extracts an interactive request voice TTS text and a voice sample in the voice request, and transmits the interactive request voice TTS text and the voice sample to the picture book interactive AI engine front-end SDK.
[0076] The voice terminal includes a voice terminal, a carrier, and a voice recognition SDK. The voice terminal receives a user's voice request through a conventional method, and transmits the voice request to the carrier. After noise reduction processing by the carrier, the voice request is transmitted to the voice recognition SDK. The voice recognition SDK extracts an interactive request voice TTS text and a voice sample in the voice request, and transcribes the interactive request voice TTS text and the voice sample to the picture book interactive AI engine front-end SDK.
[0077] The carrier is a television (set-top box), including a smart television and a set-top box, but not limited to a smart television and a set-top box, and other devices such as a projector having a similar television function.
[0078] The voice recognition SDK includes a software development kit (SDK) for implementing voice recognition of a smart television or a set-top box, and is used to provide voice recognition capability of a smart television or a set-top box system through a software development kit (SDK) to other third-party application software that does not have the capability.
[0079] The voice recognition SDK provides a voice recognition SDK interface, and the picture book application software communicates with the voice recognition SDK through the voice recognition SDK interface. The picture book application software collects voice data recognized by the voice recognition SDK through the voice recognition SDK interface and method for further processing, and pushes voice generated by the picture book application software to the voice recognition SDK and plays it to the user through the voice recognition SDK interface and method, thereby realizing voice interaction between the picture book application software and the user of the television (set-top box).
[0080] The voice recognition SDK also has an application program interface (API), and the picture book application software obtains the interactive request voice TTS text and the voice sample through the application program interface (API).
[0081] The picture book application software is an electronic picture book application program running on a smart television or a set-top box, and is installed, updated, and uninstalled in the form of APK. The picture book application software extracts and processes the interactive request voice TTS text and the voice sample of the user voice recognized by the voice recognition SDK and submits them to the picture book interactive AI engine front-end SDK. The picture book application software submits the picture book interactive response instruction audio text or the picture book interactive generated instruction audio text fed back by the picture book reading mode or the role-playing mode to the voice recognition SDK to realize voice interaction with the user. The picture book application software receives picture book metadata from the interactive picture book arrangement system and displays, arranges, and presents the picture book metadata.
[0082] The picture book interaction AI engine front-end SDK: a front-end SDK component package for providing picture book interaction AI capabilities inside a picture book application software, which contains picture book interaction AI capability API interfaces. The picture book interaction AI engine front-end SDK realizes communication coordination with the picture book interaction AI engine back-end service system through the picture book interaction AI capability API interfaces, realizes artificial intelligence analysis and processing of user interaction requests, and feeds back the analysis and processing results to the picture book interaction AI engine front-end SDK.
[0083] The picture book interaction AI engine front-end SDK realizes communication coordination with the role-playing mode program through internal calling of the picture book application software, and realizes switching of the user from the reading mode to the role-playing mode.
[0084] The interactive picture book media asset library system is a non-artificial intelligence generated ordinary interactive electronic picture book media resource metadata system, which is provided with a picture book media asset library API interface. The interactive picture book media asset library system realizes communication coordination with the picture book interaction AI engine front-end SDK through the picture book media asset library API interface, and realizes providing picture book data calling services for the picture book interaction AI engine front-end SDK in the reading mode. The interactive picture book media asset library system synchronizes picture book data through the picture book media asset library API interface and the large model pre-trained by the fairy tale corpus, and the synchronized picture book data includes but is not limited to picture book name, picture book outline, picture book script, picture book scene, picture book role, picture book action, picture book dialogue, etc. to provide the large model training and establish the picture book knowledge base. The interactive picture book media asset library system synchronizes picture book data through the picture book media asset library API interface and the picture book interaction AI engine back-end service system, and the synchronized picture book data includes but is not limited to picture book name, picture book outline, picture book script, picture book scene, picture book role, picture book action, picture book dialogue, etc. to provide the picture book interaction AI engine back-end service system to establish the picture book knowledge base. The interactive picture book media asset library system synchronizes picture book data through the picture book media asset library API interface and the interactive picture book arrangement system, and the synchronized picture book data includes but is not limited to picture book name, picture book outline, picture book script, picture book scene, picture book role, picture book action, picture book dialogue, etc. and user data such as user information and the unique identification code of the television or set-top box device currently logged in by the user, so that the interactive picture book arrangement system pushes the picture book data requested by the user to the picture book application software through the interactive picture book arrangement API interface.
[0085] The large model pre-trained by the fairy tale corpus provides an interactive picture book generation large model API interface. The role-playing mode program realizes communication coordination with the large model pre-trained by the fairy tale corpus through the interactive picture book generation large model API interface. The role-playing mode program initiates an artificial intelligence generated picture book request through the interactive picture book generation large model API interface and the large model pre-trained by the fairy tale corpus, and realizes communication coordination with the large model pre-trained by the fairy tale corpus.
[0086] The large model pre-trained by the fairy tale corpus is provided with an inter-model API interface, the large model pre-trained by the fairy tale corpus inputs the image element text of the new picture book story text generated by the large model into the text-to-image model fine-tuned by the children's picture book dataset through the inter-model API interface to generate picture book images according to defined rules; the large model pre-trained by the fairy tale corpus inputs the image element text of the new picture book story text generated by the large model into the text-to-audio model fine-tuned by the children's picture book dataset through the inter-model API interface to generate picture book audios according to defined rules.
[0087] The interactive picture book arrangement system provides an interactive picture book arrangement API interface, and the interactive picture book arrangement system communicates and cooperates with the large model pre-trained by the fairy tale corpus, the text-to-image model fine-tuned by the children's picture book dataset, the text-to-audio model fine-tuned by the children's picture book dataset, the interactive picture book media asset library system and the picture book application software through the interactive picture book arrangement API interface; the large model pre-trained by the fairy tale corpus inputs the new picture book story text generated by the large model into the interactive picture book arrangement system through the interactive picture book arrangement API interface, so that the interactive picture book arrangement system generates a picture book according to defined rules; the text-to-image model fine-tuned by the children's picture book dataset inputs the picture book images generated by the model into the interactive picture book arrangement system through the interactive picture book arrangement API interface, so that the interactive picture book arrangement system generates a picture book according to defined rules; the text-to-audio model fine-tuned by the children's picture book dataset inputs the picture book audios generated by the model into the interactive picture book arrangement system through the interactive picture book arrangement API interface, so that the interactive picture book arrangement system generates a picture book according to defined rules; the interactive picture book arrangement system pushes the picture book metadata generated according to defined rules to the picture book application software through the interactive picture book arrangement API interface.
[0088] The generation rule of the picture book generated by the artificial intelligence of the interactive picture book arrangement system is:
[0089] The picture book story storyboard images V and the picture book story storyboard audios W are automatically arranged and generated into picture book data that can be directly read by the picture book application software according to the picture book story script number S, the picture book story scene number T and the picture book story storyboard number U;
[0090] User information, a unique identification code of a television or a set-top box device currently logged in by the user is loaded on the generated picture book data, and the interactive picture book arrangement API interface and the picture book application software are communicated and cooperated.
[0091] As shown in Figures 1-3 , the present application also provides a picture book interaction method, comprising receiving a picture book reading instruction input by a user, and entering a reading mode;
[0092] In reading mode, a voice request is received and recognized, and a specific picture book is opened based on the picture book reading interaction method, or a role-playing mode is entered;
[0093] In role-playing mode, the picture book metadata of a new picture book is generated based on the user's request role through the picture book content artificial intelligence generation method, the picture book metadata is read, and the user is kept in voice interaction to enter the role-playing picture book.
[0094] Method for receiving picture book reading instructions input by a user:
[0095] The voice terminal receives the user's voice and transmits it to the carrier, and after noise reduction processing by the carrier, it is transmitted to the voice recognition SDK. The voice recognition SDK extracts the interactive request voice TTS text and the voice sample in the voice request, and transmits the interactive request voice TTS text and the voice sample to the picture book application software.
[0096] Picture book reading interaction method: When the user starts the picture book application software through the intelligent television voice remote control or other voice terminal, the default reading mode is entered, the home page is guided by voice, and the voice guide prompts the user to interact to open a specific picture book. The voice guide prompts every 20 seconds, and when the user enters the picture book reading, the voice guide prompts stop. The specific steps are as follows:
[0097] 1) The picture book application software receives the picture book reading instructions input by the user, and the TTS text and voice sample are submitted to the picture book interaction AI engine backend service system through the picture book interaction AI capability API interface by the picture book interaction AI engine frontend SDK;
[0098] The picture book interaction AI capability API interface submission should include the following information:
[0099] User information (usually mobile phone number): A+ User's current login TV or set-top box device unique identification code (usually TV or set-top box network card MAC address): B+ Voice TTS text and voice sample: C.
[0100] A takes the value of all digits of the mobile phone number;
[0101] B takes the value of all characters of the TV or set-top box network card MAC address;
[0102] C takes the value of the voice TTS text and the voice sample.
[0103] 2) The picture book interaction AI engine backend service system analyzes and processes the TTS text and voice sample, and returns the processed data to the picture book interaction AI engine frontend SDK;
[0104] The processed data should include the following key information:
[0105] User information (usually mobile phone number): A + User's current login TV or set-top box device unique identification code (usually TV or set-top box network card MAC address): B + Voiceprint recognition feature code: D + Picture book asset coding: E + Recommended picture book asset coding: F + Picture book interaction AI engine reply text: G + Analyzed interactive operation instructions: H.
[0106] D takes the value of a character representing age: 1 represents children aged 0-6, 2 represents children aged 7-12, and 3 represents adults aged 13 and above;
[0107] E takes the value of all characters of one or more picture book asset coding;
[0108] F takes the value of all characters of one or more picture book asset coding;
[0109] G takes the value of all characters of the picture book interaction AI engine reply text;
[0110] H takes the value of interactive operation instructions, specifically the following table:
[0111] Value code Code meaning 1 Flip down one page 2 Flip up one page 3 Return to the cover of the picture book 4 Flip to a user-specified page 5 Play the animation of the current page 6 Play the background music of the current page 7 Switch the background music of the current page 8 Stop playing the animation 9 Stop playing the background music 10 Mute 11 Read the text of the current page of the picture book 12 Start the game of the current picture book 13 Exit the game of the current picture book 14 Enter the role-playing mode of the current picture book
[0112] 3) The picture book interaction AI engine front-end SDK initiates a picture book data calling request to the interactive picture book media asset library system according to the received data;
[0113] The picture book interaction AI engine front-end SDK initiates a picture book data calling request to the interactive picture book media asset library system through the picture book media asset library API interface, and the request contains the following key information:
[0114] User information (usually mobile phone number): A + User's current login TV or set-top box device unique identification code (usually TV or set-top box network card MAC address): B + Picture book asset coding: D + Recommended picture book asset coding: E.
[0115] 4) The picture book media asset library system submits picture book metadata to the interactive picture book arrangement system according to the picture book data calling request;
[0116] 5) The interactive picture book arrangement system returns the picture book metadata to the picture book interaction AI engine front-end SDK in the picture book application software;
[0117] The interactive picture book arrangement system has an interactive picture book arrangement API interface, and the interactive picture book arrangement system returns the picture book metadata to the picture book interaction AI engine front-end SDK through the interactive picture book arrangement API interface;
[0118] The interactive picture book arrangement API interface contains the following key information:
[0119] User information (usually mobile phone number): A + User's current login TV or set-top box device unique identifier (usually TV or set-top box network card MAC address): B + Picture book metadata address: I + Recommended picture book metadata address: J + Picture book arrangement information (including layout order, classification, etc.): K.
[0120] 6) The picture book interaction AI engine front-end SDK displays the page based on the picture book application software according to the specific data, and submits the reply text based on the picture book interaction AI engine to the speech recognition SDK to convert it into a voice response user.
[0121] The specific data includes B + picture book metadata address I + recommended picture book metadata address J + picture book arrangement information (including layout order, classification, etc.) K.
[0122] The picture book arrangement information (including layout order, classification, etc.) K.
[0123] The user enters the picture book reading, and the interactive operation is implemented according to the value of the parsed interactive operation instruction H; the value of H includes but is not limited to Arabic numerals, English characters, Greek letters, and other legal computer characters, and Arabic numerals are used in this paper.
[0124] The voice guide prompts the user to perform the interactive operation category, which includes:
[0125] 1. Open the picture book with a specified name or screen position:
[0126] 2. Feature keywords are used to search for picture books by voice;
[0127] 3. Interactive operation in picture book reading.
[0128] When the user requests to enter the role-playing mode in the picture book reading mode by voice, the picture book application software will load the role-playing mode program through the picture book interaction AI engine front-end SDK;
[0129] When the user requests to enter the role-playing mode in the picture book reading mode by voice through the smart TV voice remote control or other voice terminal, the picture book application software will load the role-playing mode program through the picture book interaction AI engine front-end SDK, at which time the role-playing mode program will perform voice guidance, and the role voice guidance prompts the user to speak the selected picture book role. The voice guidance is played once every 20 seconds, and the voice guidance stops when the user enters the role-playing picture book.
[0130] As shown in Figure 4 , the present application also provides a picture book content artificial intelligence generation method, which comprises the following steps:
[0131] 1) The user selects a picture book role by voice under the voice guidance of the role-playing mode program;
[0132] 2) The picture book application software obtains the interactive request voice TTS text and the voice sample through the voice recognition SDK and submits the TTS text and the voice sample to the picture book interactive AI engine front-end SDK;
[0133] 3) The picture book interactive AI engine front-end SDK submits the TTS text and the voice sample to the picture book interactive AI engine back-end service system through the picture book interactive AI capability API interface,
[0134] The picture book interactive AI capability API interface submission should include the following information:
[0135] User information (usually mobile phone number): A + unique identification code of the television or set-top box device currently logged in by the user (usually the television or set-top box network card MAC address): B
[0136] + voice TTS text and the voice sample: C.
[0137] A takes the value of all digits of the mobile phone number;
[0138] B takes the value of all characters of the television or set-top box network card MAC address;
[0139] C takes the value of the voice TTS text and the voice sample.
[0140] 4) The picture book interactive AI engine back-end service system analyzes and processes the TTS text and the voice sample and returns the processed data to the picture book interactive AI engine front-end SDK;
[0141] The processed data includes the following key information:
[0142] User information (usually mobile phone number): A + unique identification code of the television or set-top box device currently logged in by the user (usually the television or set-top box network card MAC address): B
[0143] + voiceprint recognition feature code: D + reply text based on the picture book interactive AI engine:
[0144] G + parsed interactive operation instructions: H + parsed user role selection information: L.
[0145] D takes the value of characters representing age: 1 represents 0-6 year-old children, 2 represents 7-12 year-old children, and 3 represents adults over 13 years old;
[0146] G takes the value of all characters of the reply text of the picture book interactive AI engine;
[0147] H takes the value of the interactive operation instructions, which are as follows:
[0148]
[0149]
[0150] L takes all the characters of the user role selection information parsed by the picture book interaction AI engine;
[0151] 5) The picture book interaction AI engine front-end SDK forwards the received voiceprint recognition, picture book interaction response processing data, and user role selection information to the role-playing mode program, which forwards the received data to the large model pre-trained by the fairy tale corpus.
[0152] 6) The large model pre-trained by the fairy tale corpus analyzes and processes the received data to generate three pieces of text data: ① new picture book story text generated by the large model, ② image element text of the text generated by the large model, and ③ audio element text of the text generated by the large model.
[0153] 7) The new picture book story text generated by the large model pre-trained by the fairy tale corpus is output to the interactive picture book arrangement system.
[0154] The generated new picture book story text contains the following key information classification generation:
[0155] Picture book story outline: M + picture book story script: N + picture book story scene: O + picture book role: P + picture book scene and role action: Q + role dialogue: R.
[0156] Picture book story outline M takes the summary text of picture book story script N according to the order of story development;
[0157] Picture book story script N takes the description text of picture book story scene O according to the order of story development;
[0158] Picture book story scene O takes the picture book story scene description text (scene description text is not limited to picture background);
[0159] Picture book role P takes the role text contained in story scene O;
[0160] Picture book scene and role action Q takes the scene and role action text contained in story scene O;
[0161] Role dialogue R takes the role dialogue text contained in story scene O.
[0162] 8) The image element text of the text generated by the large model pre-trained by the fairy tale corpus is output to the text-to-image model fine-tuned by the children's picture book dataset.
[0163] The generated picture book text image element text contains the following key information classification generation:
[0164] Picture book story scene: O + Picture book character: P + Picture book scene and character action: Q + Character dialogue: R + Picture book story script serial number: S + Picture book story scene serial number: T + Picture book story shot serial number: U.
[0165] Wherein the picture book story scene O, the picture book character P, the picture book scene and character action Q, the character dialogue R take values and the foregoing definitions are consistent,
[0166] Wherein the picture book story script serial number R is the serial number of the picture book story outline M in the order of story development;
[0167] Wherein the picture book story scene serial number T is the serial number of the picture book story outline N in the order of story development;
[0168] Wherein the picture book story shot serial number U is the combination of the script serial number R and the scene serial number T.
[0169] Wherein the audio element text of the generated picture book text contains the following key information classification generation:
[0170] Character dialogue: R + Picture book story script serial number: S + Picture book story scene serial number: T + Picture book story shot serial number: U,
[0171] Wherein the character dialogue R, the picture book story script serial number S, the picture book story scene serial number T take values, and the picture book story shot serial number U take values and the foregoing definitions are consistent.
[0172] 10) The text-to-image model fine-tuned by the children's picture book dataset generates picture book shot images according to the generated text image element text,
[0173] Wherein the generated picture book shot image should contain the following key information classification generation:
[0174] Picture book story script serial number: S + Picture book story scene serial number: T + Picture book story shot image serial number: U + Picture book story shot image: V,
[0175] The picture book story script serial number S, the picture book story scene serial number T, and the picture book story shot serial number U take values and the foregoing definitions are consistent,
[0176] The picture book story shot image V takes the value of the image generated by the text-to-image model.
[0177] 11) The text-to-audio model fine-tuned by the children's picture book dataset generates picture book shot audio according to the generated text audio element text,
[0178] Wherein the generated picture book shot image should contain the following key information classification generation:
[0179] The picture book story script serial number S, the picture book story scene serial number T, and the picture book story frame serial number U take values consistent with the foregoing definitions,
[0180] The picture book story script serial number S, the picture book story scene serial number T, and the picture book story frame serial number U take values consistent with the foregoing definitions,
[0181] The picture book story frame audio W takes a value of audio generated by a text-to-speech audio model.
[0182] 12) The interactive picture book arrangement system automatically arranges the picture book story frame images V and the picture book story frame audio W according to the received generated content according to the picture book story script serial number S, the picture book story scene serial number T, and the picture book story frame serial number U to generate picture book metadata that can be directly read by a picture book application software.
[0183] 13) The picture book application software reads the automatically arranged picture book metadata, simultaneously responds to a user's voice interaction request through a picture book interaction AI engine front-end SDK, and maintains interaction with the user through a voice recognition SDK in audio, to realize a picture book and interaction automatically generated by artificial intelligence.
[0184] Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor should belong to the scope of protection of the present application.
Claims
1. A picture book interaction system, characterized by: The application comprises a picture book application software and a server system, the picture book application software comprises a picture book interactive AI engine front-end SDK and a role-playing mode program, and the server system comprises a picture book interactive AI engine back-end service system, an interactive picture book media asset library system and a picture book content artificial intelligence generation system, The picture book interactive AI engine front-end SDK is used for realizing picture book interaction, picture book media asset data calling and picture book content artificial intelligence generation system generation of a picture book; The role-playing mode program: in the reading mode, according to the voice request of a user, the picture book interactive AI engine front-end SDK is internally called and initiates a picture book request to the picture book content artificial intelligence generation system; The picture book interactive AI engine back-end service system: provides cloud AI service capability for the picture book interactive AI engine front-end SDK, provides artificial intelligence analysis and processing of user interactive requests, and feeds back the analysis and processing results to the picture book interactive AI engine front-end SDK; The interactive picture book media asset library system: is an interactive electronic picture book media resource metadata system, contains picture book image data and audio data, and provides picture book data calling service for the picture book interactive AI engine front-end SDK; The picture book content artificial intelligence generation system: artificially generates picture book metadata of a new picture book based on a user request role, and pushes the picture book metadata to the picture book application software, which receives and displays the picture book metadata.
2. The picture book interaction system according to claim 1, characterized in that: The picture book content artificial intelligence generation system comprises a large model pre-trained through a fairy tale corpus, a text-to-image model fine-tuned through a children's picture book data set, a text-to-audio model fine-tuned through the children's picture book data set and an interactive picture book arrangement system, the large model pre-trained through the fairy tale corpus artificially generates a new picture book story text based on a user request role, the text-to-image model fine-tuned through the children's picture book data set inputs and outputs new picture book images from the new picture book story text generated by the large model; the text-to-audio model fine-tuned through the children's picture book data set inputs and outputs new picture book audio from the new picture book story text generated by the large model; the interactive picture book arrangement system combines the picture book images and the picture book audio to generate a new picture book story, automatically combines the new picture book story into new picture book metadata and pushes the picture book metadata to the picture book application software.
3. The picture book interaction system according to claim 2, characterized in that: The picture book interactive AI engine front-end SDK realizes voice interaction with a voice terminal, the voice terminal receives a voice request of a user and analyzes and processes the voice request, and transmits the processed data to the picture book application software.
4. The picture book interaction system according to claim 3, characterized in that: The voice terminal comprises a voice terminal, a carrier and a voice recognition SDK, the voice terminal receives a voice request of a user, transmits the voice request to the carrier, transmits the voice request to the voice recognition SDK after noise reduction processing by the carrier, the voice recognition SDK extracts an interactive request voice TTS text and a voice sample in the voice request, and transmits the interactive request voice TTS text and the voice sample to the picture book interactive AI engine front-end SDK.
5. The picture book interactive system of claim 3, wherein: The picture book interaction AI engine front-end SDK includes a picture book interaction AI capability API interface. The picture book interaction AI engine front-end SDK communicates and cooperates with the picture book interaction AI engine back-end service system through the picture book interaction AI capability API interface, and analyzes and processes an artificial intelligence of a user interaction request, and feeds back analysis and processing results to the picture book interaction AI engine front-end SDK.
6. A picture book interaction method, characterized by: Comprise Receiving a picture book reading instruction input by a user, entering a reading mode; In the reading mode, receiving a voice request and identifying the voice request, and opening a specific picture book based on a picture book reading interaction method, or entering a role-playing mode; In the role-playing mode, generating picture book metadata of a new picture book based on a user-requested role through a picture book content artificial intelligence generation method, reading the picture book metadata, and maintaining voice interaction with the user to make the user enter a role-playing picture book.
7. The method of claim 6, wherein: The picture book reading interaction method comprises the following steps: 1) The picture book application software receives a picture book reading instruction input by a user. The picture book interaction AI engine front-end SDK submits TTS text and a voice sample to the picture book interaction AI engine back-end service system through a picture book interaction AI capability API interface; 2) The picture book interaction AI engine back-end service system analyzes and processes the TTS text and the voice sample, and returns processed data to the picture book interaction AI engine front-end SDK; 3) The picture book interaction AI engine front-end SDK initiates a picture book data calling request to an interactive picture book media asset library system according to the received data; 4) The picture book media asset library system submits picture book metadata to an interactive picture book arrangement system according to the picture book data calling request; 5) The interactive picture book arrangement system returns the picture book metadata to the picture book interaction AI engine front-end SDK in the picture book application software; 6) The picture book interaction AI engine front-end SDK performs page display based on the picture book application software according to specific data, and submits a reply text based on the picture book interaction AI engine to a speech recognition SDK to convert the reply text into a voice response to the user.
8. The method of claim 7, wherein: The method of receiving a picture book reading instruction input by a user comprises the following steps:
9. A picture book content artificial intelligence generation method, characterized by: 1) The user selects a picture book role through voice under the voice guidance prompt of the role-playing mode program; 2) The picture book application software acquires interactive request voice TTS text and a voice sample through a speech recognition SDK, and submits the TTS text and the voice sample to the picture book interaction AI engine front-end SDK; 3) The picture book interaction AI engine front-end SDK submits the TTS text and the voice sample to the picture book interaction AI engine back-end service system through a picture book interaction AI capability API interface; 4) The picture book interaction AI engine back-end service system analyzes and processes the TTS text and the voice sample, and returns processed data to the picture book interaction AI engine front-end SDK; 5) The e-book interaction AI engine front-end SDK forwards the received voiceprint recognition, e-book interaction response processing data, and user role selection information to the role-playing mode program, which forwards the received data to the large model pre-trained with the fairy tale corpus; 6) The large model pre-trained with the fairy tale corpus analyzes and processes the received data to generate three pieces of text data: ① new e-book story text generated by the large model, ② image element text of the text generated by the large model, and ③ audio element text of the text generated by the large model; 7) The new e-book story text generated by the large model pre-trained with the fairy tale corpus is output to the interactive e-book arrangement system; 8) The image element text of the text generated by the large model pre-trained with the fairy tale corpus is output to the text-to-image model fine-tuned with the children's e-book dataset; 9) The audio element text of the text generated by the large model pre-trained with the fairy tale corpus is output to the text-to-audio model fine-tuned with the children's e-book dataset; 10) The text-to-image model fine-tuned with the children's e-book dataset generates e-book split image according to the generated image element text of the text; 11) The text-to-audio model fine-tuned with the children's e-book dataset generates e-book split audio according to the generated audio element text of the text; 12) The interactive e-book arrangement system automatically arranges and generates e-book metadata that can be read by e-book application software according to the received generated content, e-book story script number S, e-book story scene number T, and e-book story split number U, and generates e-book story split image V and e-book story split audio W; 13) The e-book application software reads the automatically arranged and generated e-book metadata, and simultaneously responds to the user's voice interaction request through the e-book interaction AI engine front-end SDK and maintains interaction with the user through voice recognition SDK in audio.
10. The picture book content artificial intelligence generation method of claim 9, wherein: The generated new e-book story text contains the following key information classification generation: E-book story outline: M + e-book story script: N + e-book story scene: O + e-book character: P + e-book scene and character action: Q + character dialogue: R; E-book story outline M is the summary text of e-book story script N in the order of story development; E-book story script N is the description text of e-book story scene O in the order of story development; E-book story scene O is the description text of e-book story scene; E-book character P is the character text contained in story scene O; E-book scene and character action Q is the scene and character action text contained in story scene O; Character dialogue R is the character dialogue text contained in story scene O.
Citation Information
Patent Citations
Television interaction method and system based on application scene large model intelligent decision
CN119402690A
Cited By
AI emotion interaction guiding system for children picture book reading
CN121807207A