Picture book rewriting method, picture book sequel writing method, and display device

By storing multi-dimensional information in picture books and using various processing models for automated processing, the problem of cumbersome traditional picture book rewriting process is solved, achieving efficient picture book rewriting and continuation while maintaining the consistency of the picture book.

CN119991873BActive Publication Date: 2025-11-21HISENSE VISUAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411982780.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-11-21
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

The traditional process of rewriting picture books is cumbersome, resulting in low efficiency.

Method used

By associating picture book identifiers with multi-dimensional information and using text processing, image processing, and audio processing models for automated processing, the content of picture books can be automatically rewritten and continued, maintaining the consistency of the characters, style, and theme of the picture books.

Benefits of technology

It reduces manual steps, improves the efficiency of rewriting and continuing picture books, and ensures that the rewritten content is consistent with the original picture book in terms of characters, style, theme, etc.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991873B_ABST
    Figure CN119991873B_ABST
Patent Text Reader

Abstract

The application relates to a picture book rewriting method, a picture book sequel writing method and a display device, and relates to the technical field of display devices. The method comprises the following steps: receiving rewriting requirement information, taking the identification of a current picture book as picture book identification; identifying text information and acquiring picture book information from a database according to the picture book identification; inputting information representing a role into a text processing model to obtain rewritten text content; inputting the rewritten text content, information representing a style and information representing a role feature into an image processing model to obtain rewritten image data; performing playing audio data synthesis processing to obtain broadcast audio data, and performing background audio data generation processing to obtain background audio data; replacing picture book image data with the rewritten image data, and replacing broadcast audio data and background audio data with broadcast audio data and background audio data corresponding to the rewritten image data to obtain a rewritten picture book. The method can improve the rewriting efficiency of the picture book.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of display devices, and in particular to a picture book rewriting method, a picture book continuation method and a display device. BACKGROUND

[0002] The display device such as a smart television refers to a device that realizes a bidirectional man-machine interaction function based on an Internet application technology.

[0003] In the prior art, when a generated picture book does not meet the user's expectation and the user wants to rewrite the picture book, a complete picture book is usually generated again according to the user's new requirement. However, the picture book generation process needs to be executed again every time the picture book is rewritten, which leads to a cumbersome picture book rewriting process and low picture book rewriting efficiency.

[0004] Therefore, the prior art has the technical problem of low picture book rewriting efficiency. SUMMARY

[0005] The present application provides a picture book rewriting method, a picture book continuation method and a display device to solve the technical problem of low picture book rewriting efficiency.

[0006] In a first aspect, some embodiments provide a picture book rewriting method applied to a display device, and the method comprises:

[0007] In a case where a picture book content is displayed on a user interface, receiving rewriting requirement information for the picture book content, and taking an identifier of a current picture book as a picture book identifier corresponding to the rewriting requirement information;

[0008] Identifying text information corresponding to the rewriting requirement information, and acquiring picture book information corresponding to the picture book identifier from a database according to the picture book identifier; the picture book information represents multi-dimensional information of a picture book corresponding to the picture book identifier;

[0009] Inputting the text information and information representing a role in the picture book information into a text processing model to obtain rewritten text content; the rewritten text content represents text content composed of the text information and the information representing the role; the text processing model is used to perform picture book text content generation processing according to the input text information and the information representing the role, and output the rewritten text content;

[0010] input the rewritten text content, information representing style in the picture book information, and information representing character features in the picture book information into an image processing model to obtain rewritten image data; the rewritten image data represents image data that matches the rewritten text content, the information representing style, and the information representing character features; the image processing model is used to perform picture book image data generation processing according to the input rewritten text content, the information representing style, and the information representing character features, and output the rewritten image data;

[0011] based on the rewritten text content and information representing a broadcast tone in the picture book information, perform broadcast audio data synthesis processing to obtain broadcast audio data corresponding to the rewritten image data, and based on information representing a theme in the picture book information and the information representing style, perform background audio data generation processing to obtain background audio data corresponding to the rewritten image data;

[0012] replace picture book image data in the picture book content in the picture book corresponding to the picture book identifier with the rewritten image data, and replace broadcast audio data and background audio data corresponding to the picture book image data with broadcast audio data and background audio data corresponding to the rewritten image data, to obtain a rewritten picture book; the number of pages of the rewritten picture book remains unchanged from the number of pages of the picture book corresponding to the picture book identifier.

[0013] Technical effects: By storing the picture book identifier in association with multi-dimensional information in the database, and uniformly obtaining information such as characters, styles, and themes based on the picture book identifier during rewriting, it is beneficial to avoid redundant operations such as repeated information acquisition and manual input. By inputting the obtained information into a text processing model, an image processing model, and an audio processing model for automatic processing, and accurately replacing the information while keeping the number of pages unchanged, it is beneficial to maintain the consistency of the rewritten content and the original picture book in terms of characters, styles, themes, and the like. Thus, it is beneficial to realize automatic rewriting of picture books, reduce manual operation steps, and improve the rewriting efficiency of picture books.

[0014] In some embodiments of the present application, the receiving of the rewriting requirement information for the picture book content in the case of displaying the picture book content in the user interface includes:

[0015] In the case of displaying the picture book content in the user interface, receiving an input picture book content selection operation;

[0016] In the case of identifying the selected picture book content according to the picture book content selection operation, receiving the rewriting requirement information for the picture book content.

[0017] Technical effects: By receiving the user's selection operation on the displayed picture book content first, and then identifying the specific picture book content based on the selection operation, and then receiving the rewriting demand information for the selected content, it is beneficial to accurately locate the specific content position that the user wants to rewrite, avoiding the problem of unclear rewriting range. Therefore, it is beneficial to ensure that there is a clear corresponding relationship between the rewriting demand information and the picture book content selected by the user, and to improve the accuracy of picture book rewriting.

[0018] In some embodiments of the present application, after receiving the input picture book content selection operation, further comprising:

[0019] In response to the picture book content selection operation, identifying the picture book content corresponding to the current focus;

[0020] Identifying the picture book content corresponding to the current focus as the selected picture book content.

[0021] Technical effects: By responding to the picture book content selection operation and identifying the picture book content corresponding to the current focus, it is beneficial to accurately locate the specific content position that the user wants to rewrite, avoiding incorrect identification of the user's rewriting intention. By identifying the picture book content corresponding to the current focus as the selected picture book content, it is beneficial to establish an accurate corresponding relationship between the focus position and the specific content. Therefore, it is beneficial to accurately understand the user's rewriting demand and improve the accuracy of picture book rewriting.

[0022] In some embodiments of the present application, before obtaining the picture book information corresponding to the picture book identifier from the database according to the picture book identifier corresponding to the rewriting demand information, further comprising:

[0023] Performing intent recognition processing on the text information corresponding to the rewriting demand information to obtain an intent recognition result corresponding to the rewriting demand information; the intent recognition result is used to represent whether to obtain data from the database;

[0024] Based on the intent recognition result, determining whether to obtain data from the database;

[0025] The picture book information corresponding to the picture book identifier is obtained from the database according to the picture book identifier corresponding to the rewriting demand information, comprising:

[0026] In the case of determining to obtain data from the database, obtaining the picture book information corresponding to the picture book identifier from the database according to the picture book identifier corresponding to the rewriting demand information.

[0027] Technical effects: By performing intent recognition processing on the rewritten demand information corresponding text information, and judging whether data needs to be obtained from the database based on the intent recognition result, it is beneficial to avoid unnecessary database access operations; by obtaining the picture book information from the database only when it is indeed needed, it is beneficial to implement on-demand execution of data acquisition operations; thereby it is beneficial to reduce resource consumption while ensuring normal implementation of rewriting functions, and improve overall operation efficiency.

[0028] In some embodiments of the present application, the information representing the role characteristics is also obtained in the following way:

[0029] From the picture book corresponding to the picture book identifier, picture book image data containing information representing the role is selected as reference picture book image data;

[0030] From the reference picture book image data, information representing the role characteristics is extracted.

[0031] Technical effects: By selecting picture book image data containing information representing the role from the picture book as a reference, it is beneficial to obtain standardized visual features of the role; by extracting information representing the role characteristics from the reference picture book image data, it is beneficial to establish baseline data of the role characteristics; thereby it is beneficial to maintain consistency of the role image in the subsequent picture book generation process, and improve the quality of picture book generation.

[0032] In some embodiments of the present application, the inputting the text information and the information representing the role in the picture book information into the text processing model to obtain the rewritten text content includes:

[0033] Input the text information into the text processing model to obtain key text information; the key text information represents the key information in the text information; the text processing model is used to output the key text information in the text information according to the input text information;

[0034] Input the key text information and the information representing the role into the text processing model to obtain the rewritten text content; the rewritten text content represents the text content composed of the key text information and the information representing the role; the text processing model is used to perform picture book text content generation processing according to the input key text information and the information representing the role, and output the rewritten text content.

[0035] Technical effects: By inputting the text information into the text processing model to extract key text information, it is beneficial to accurately grasp the core content of the user's rewriting demand; by inputting the extracted key text information and the information representing the role into the text processing model, it is beneficial to generate new text content while maintaining the characteristics of the role; thereby it is beneficial to ensure that the rewritten content meets the user's demand while maintaining the coherence of the picture book role image, improving the generation quality of the rewritten text content.

[0036] In some embodiments of the present application, the inputting the rewritten text content, the information representing the style in the picture book information, and the information representing the characteristics of the role in the picture book information into the image processing model to obtain rewritten image data comprises:

[0037] inputting the rewritten text content into the text processing model to obtain image generation prompt information corresponding to the rewritten text content; the image generation prompt information is used to represent text content adapted to picture book image data generation processing; the text processing model is used to adjust the input rewritten text content and output the adjusted text content; the adjusted text content is the image generation prompt information;

[0038] inputting the image generation prompt information, the information representing the style, and the information representing the characteristics of the role into an image processing model to obtain rewritten image data; the rewritten image data represents image data matched with the image generation prompt information, the information representing the style, and the information representing the characteristics of the role; the image processing model is used to perform picture book image data generation processing according to the input image generation prompt information, the information representing the style, and the information representing the characteristics of the role, and output the rewritten image data.

[0039] Technical effects: By converting the rewritten text content into more suitable prompt information for image generation, it is beneficial to improve the accuracy of text-to-image conversion; by inputting the image generation prompt information, the information representing the style, and the information representing the characteristics of the role into the image processing model, it is beneficial to maintain the consistency of the style and the characteristics of the role during image generation; thereby it is beneficial to generate high-quality image data that meets the user's rewriting demand and maintains the original picture book style and role characteristics.

[0040] In some embodiments of the present application, the inputting the image generation prompt information, the information representing the style, and the information representing the characteristics of the role into the image processing model to obtain rewritten image data comprises:

[0041] inputting the image generation prompt information and the information representing the character features into an image processing model to obtain initial rewriting image data; the initial rewriting image data represents image data matched with the image generation prompt information and the information representing the character features; the image processing model is used to perform picture book image data generation processing according to the input image generation prompt information and the information representing the character features, and output the initial rewriting image data;

[0042] inputting the initial rewriting image data and the information representing the style into the image processing model to obtain rewriting image data; the style information of the rewriting image data is the same as the information representing the style, and the image content of the rewriting image data is the same as that of the initial rewriting image data; the image processing model is used to perform rendering processing of the information representing the style on the input initial rewriting image data, and output the rewriting image data.

[0043] Technical effects: By dividing the generation process of rewriting image data into two steps, first generating initial rewriting image data according to image generation prompt information and information representing character features, which is conducive to ensuring the content accuracy of the generated image and the consistency of the character features; and then performing style rendering processing on the initial rewriting image data to obtain rewriting image data, which is conducive to realizing the unity of style while keeping the image content unchanged; thereby improving the generation quality of rewriting image data, so that the generated image can accurately express the content and maintain the consistency of the style.

[0044] In some embodiments of the present application, based on the rewriting text content and the information representing the voice in the picture book information, performing playing audio data synthesis processing to obtain playing audio data corresponding to the rewriting image data, and based on the information representing the theme in the picture book information and the information representing the style, performing background audio data generation processing to obtain background audio data corresponding to the rewriting image data, comprising:

[0045] inputting the rewriting text content and the information representing the voice into a playing audio data processing model to obtain playing audio data corresponding to the rewriting image data; the timbre information of the playing audio data is the same as the information representing the voice, and the text content corresponding to the playing audio data is the same as the rewriting text content; the playing audio data processing model is used to convert the input rewriting text content into corresponding initial playing audio data, perform adjustment processing of the information representing the voice on the initial playing audio data, and output playing audio data corresponding to the rewriting image data;

[0046] input the information representing the theme and the information representing the style into a background audio data processing model to obtain background audio data corresponding to the rewritten image data; a background audio label of the background audio data is a background audio label corresponding to the information representing the theme and the information representing the style; the background audio data processing model is used for performing background audio data generation processing according to a background audio label corresponding to the input information representing the theme and the information representing the style, and outputting the background audio data corresponding to the rewritten image data.

[0047] Technical effects: By separating the generation processes of the broadcast audio data and the background audio data, it is beneficial to control the timbre characteristics of the broadcast content and the style characteristics of the background music respectively; by performing timbre adjustment processing on the broadcast audio data and theme and style label matching processing on the background audio data, it is beneficial to ensure the consistency of the generated audio with the rewritten content, theme and style.

[0048] In a second aspect, some embodiments also provide a picture book continuation method applied to a display device, the method comprising:

[0049] In a case where a picture book is displayed on a user interface, receiving continuation demand information for the picture book, and taking an identifier of the current picture book as a picture book identifier corresponding to the continuation demand information;

[0050] Identifying text information corresponding to the continuation demand information, and according to the picture book identifier corresponding to the continuation demand information, obtaining picture book information corresponding to the picture book identifier from a database; the picture book information represents multi-dimensional information of a picture book corresponding to the picture book identifier;

[0051] inputting the text information and information representing a character in the picture book information into a text processing model to obtain continuation text content; the continuation text content represents text content composed of the text information and the information representing the character; the text processing model is used for performing picture book text content generation processing according to the input text information and the information representing the character, and outputting the continuation text content;

[0052] inputting the continuation text content, information representing a style in the picture book information and information representing a character feature in the picture book information into an image processing model to obtain continuation image data; the continuation image data represents image data matched with the continuation text content, the information representing the style and the information representing the character feature; the image processing model is used for performing picture book image data generation processing according to the input continuation text content, the information representing the style and the information representing the character feature, and outputting the continuation image data;

[0053] Based on the continuation text content and the information representing the tone in the picture book information, play audio data synthesis processing is performed to obtain play audio data corresponding to the continuation image data, and based on the information representing the theme in the picture book information and the information representing the style, background audio data generation processing is performed to obtain background audio data corresponding to the continuation image data;

[0054] The semantic recognition processing is performed on the text information corresponding to the continuation demand information to obtain a continuation position of the continuation image data in the picture book corresponding to the picture book identifier. The continuation image data is added in the continuation position in the picture book corresponding to the picture book identifier, and the play audio data and the background audio data corresponding to the continuation image data are configured to obtain a continuation picture book. The number of pages of the continuation picture book is greater than the number of pages of the picture book corresponding to the picture book identifier.

[0055] Technical effects: By obtaining multi-dimensional information of the original picture book as a constraint condition for continuation, it is beneficial to ensure that the continuation content is consistent with the original picture book in terms of character image, artistic style and story theme. By cooperatively generating text content, image data, play audio and background audio, and determining a suitable continuation position according to semantic recognition, it is beneficial to realize natural connection of multi-modal content, to expand the content of the picture book while maintaining the coherence and integrity of the story. Thus, it is beneficial to realize automatic continuation of the picture book, reduce manual operation steps, and improve the continuation efficiency of the picture book.

[0056] In some embodiments of the present application, the receiving of the continuation demand information for the picture book in the case of displaying the picture book in the user interface comprises:

[0057] In the case of displaying the picture book in the user interface, receiving an input picture book selection operation;

[0058] In response to the picture book selection operation, identifying the picture book corresponding to the current focus;

[0059] The picture book corresponding to the current focus is identified as the selected picture book;

[0060] In the case of identifying the selected picture book according to the picture book selection operation, the continuation demand information for the picture book is received.

[0061] Technical effects: By responding to the user's picture book selection operation and identifying the picture book corresponding to the current focus position, it is beneficial to establish a clear correspondence between user operation and specific picture books; By identifying the picture book corresponding to the current focus as the picture book selected by the user, it is beneficial to accurately locate the specific picture book that the user wants to continue writing; thereby it is beneficial to improve the accuracy of the user's picture book selection operation; By receiving the continuation writing demand information after identifying the picture book selected by the user, it is beneficial to ensure the accurate correspondence between the continuation writing demand and the specific picture book; thereby it is beneficial to establish an orderly picture book continuation writing operation process, avoid incorrect matching between the continuation writing demand and the picture book, and improve the accuracy of picture book continuation writing.

[0062] In some embodiments of the present application, the adding the continuation image data in the continuation writing position in the picture book corresponding to the picture book identifier, and configuring the playback audio data and the background audio data corresponding to the continuation image data to obtain a picture book after continuation writing, comprises:

[0063] In the case where the continuation writing position is the end of the picture book corresponding to the picture book identifier, the continuation image data is added after the last page in the picture book corresponding to the picture book identifier, and the playback audio data and the background audio data corresponding to the continuation image data are configured to obtain a picture book after continuation writing;

[0064] In the case where the continuation writing position is the beginning of the picture book corresponding to the picture book identifier, the continuation image data is added before the first page in the picture book corresponding to the picture book identifier, and the playback audio data and the background audio data corresponding to the continuation image data are configured to obtain a picture book after continuation writing.

[0065] Technical effects: By flexibly adding continuation image data at the beginning or end of the picture book, it is beneficial for the user to select a suitable continuation writing position according to actual needs; By configuring the corresponding playback audio data and background audio data for the continuation image data, it is beneficial to maintain the consistency of the continuation content and the original picture book in visual and auditory experience.

[0066] The third aspect, some embodiments also provide a display device, comprising:

[0067] a display;

[0068] a controller coupled to the display and configured to:

[0069] In the case where the user interface displays picture book content, receive rewriting demand information for the picture book content, and take the identifier of the current picture book as the picture book identifier corresponding to the rewriting demand information;

[0070] identify the text information corresponding to the rewriting requirement information, and obtain picture book information corresponding to the picture book identifier according to the picture book identifier corresponding to the rewriting requirement information; the picture book information represents multi-dimensional information of a picture book corresponding to the picture book identifier;

[0071] input the text information and the information representing the role in the picture book information into a text processing model to obtain rewritten text content; the rewritten text content represents text content composed of the text information and the information representing the role; the text processing model is used for picture book text content generation processing according to the input text information and the information representing the role, and outputs the rewritten text content;

[0072] input the rewritten text content, the information representing the style in the picture book information, and the information representing the role characteristics in the picture book information into an image processing model to obtain rewritten image data; the rewritten image data represents image data matched with the rewritten text content, the information representing the style, and the information representing the role characteristics; the image processing model is used for picture book image data generation processing according to the input rewritten text content, the information representing the style, and the information representing the role characteristics, and outputs the rewritten image data;

[0073] based on the rewritten text content and the information representing the tone in the picture book information, perform a play audio data synthesis processing to obtain play audio data corresponding to the rewritten image data, and based on the information representing the theme in the picture book information and the information representing the style, perform a background audio data generation processing to obtain background audio data corresponding to the rewritten image data;

[0074] replace the picture book image data in the picture book content in the picture book corresponding to the picture book identifier with the rewritten image data, and replace the play audio data and the background audio data corresponding to the picture book image data with the play audio data and the background audio data corresponding to the rewritten image data, to obtain a rewritten picture book; the number of pages of the rewritten picture book remains unchanged.

[0075] Technical effects: by storing the picture book identifier and the multi-dimensional information in the database, and uniformly obtaining the role, style, theme and other information based on the picture book identifier during rewriting, it is beneficial to avoid redundant operations such as repeated information acquisition and manual input; by inputting the obtained information into the text processing model, the image processing model and the audio processing model for automatic processing, and accurately replacing while keeping the number of pages unchanged, it is beneficial to keep the consistency of the rewritten content and the original picture book in terms of role, style, theme and other aspects; thereby it is beneficial to realize the automatic rewriting of the picture book, reduce the manual operation steps, and thereby improve the rewriting efficiency of the picture book.

[0076] In a fourth aspect, some embodiments further provide a display device, comprising:

[0077] a display;

[0078] a controller coupled to the display and configured to:

[0079] in a case where a picture book is displayed on the user interface, receive a continuation demand information for the picture book, and take an identification of the picture book as a picture book identification corresponding to the continuation demand information;

[0080] identify text information corresponding to the continuation demand information, and acquire picture book information corresponding to the picture book identification from a database according to the picture book identification corresponding to the continuation demand information; the picture book information represents multi-dimension information of a picture book corresponding to the picture book identification;

[0081] input information representing a character in the text information and the picture book information into a text processing model to obtain a continuation text content; the continuation text content represents a text content composed of the text information and the information representing the character; the text processing model is used to perform a picture book text content generation processing according to the input text information and the information representing the character, and output the continuation text content;

[0082] input the continuation text content, information representing a style in the picture book information, and information representing a character feature in the picture book information into an image processing model to obtain continuation image data; the continuation image data represents image data matched with the continuation text content, the information representing the style, and the information representing the character feature; the image processing model is used to perform a picture book image data generation processing according to the input continuation text content, the information representing the style, and the information representing the character feature, and output the continuation image data;

[0083] perform a playing audio data synthesis processing based on the continuation text content and information representing a playing tone in the picture book information to obtain playing audio data corresponding to the continuation image data, and perform a background audio data generation processing based on information representing a theme in the picture book information and the information representing the style to obtain background audio data corresponding to the continuation image data;

[0084] perform semantic recognition processing on the text information corresponding to the continuation demand information, to obtain a continuation position of the continuation image data in the picture book corresponding to the picture book identifier; and add the continuation image data in the continuation position in the picture book corresponding to the picture book identifier, and configure the playback audio data and the background audio data corresponding to the continuation image data, to obtain a picture book after the continuation; and the number of pages of the picture book after the continuation is greater than the number of pages of the picture book corresponding to the picture book identifier.

[0085] Technical effects: By obtaining multi-dimensional information of the original picture book and taking it as a constraint condition for the continuation, it is beneficial to ensure that the continuation content is consistent with the original picture book in terms of character image, artistic style and story theme; by cooperatively generating text content, image data, playback audio and background audio, and determining a suitable continuation position according to semantic recognition, it is beneficial to realize natural connection of multi-modal content, to expand the content of the picture book while maintaining the coherence and integrity of the story; thereby, it is beneficial to realize automatic continuation of the picture book, to reduce manual operation steps, and to improve the continuation efficiency of the picture book. BRIEF DESCRIPTION OF DRAWINGS

[0086] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the following will briefly introduce the drawings needed to be used in the embodiment or related art description. Obviously, the drawings in the following description only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0087] Figure 1 The schematic diagram of the operation scene between the display device and the control device provided by some embodiments of the present application;

[0088] Figure 2 The schematic diagram of the hardware configuration of the display device provided by some embodiments of the present application;

[0089] Figure 3 The schematic diagram of the hardware configuration of the control device provided by some embodiments of the present application;

[0090] Figure 4 The schematic diagram of the software configuration of the display device provided by some embodiments of the present application;

[0091] Figure 5 The flowchart of the picture book rewriting method provided by some embodiments of the present application;

[0092] Figure 6 The flowchart of the step of receiving the rewriting demand information provided by some embodiments of the present application;

[0093] Figure 7 The image schematic diagram of the picture book content provided by some embodiments of the present application;

[0094] Figure 8 A flowchart of steps of identifying the content of a picture book for some embodiments of the present application;

[0095] Figure 9 A flowchart of a picture book continuation method for some embodiments of the present application;

[0096] Figure 10 An image diagram of a picture book for some embodiments of the present application;

[0097] Figure 11 A signaling interaction diagram of a picture book rewriting method for some embodiments of the present application;

[0098] Figure 12 An overall architecture diagram of a picture book rewriting method for some embodiments of the present application;

[0099] Figure 13 A main branch architecture diagram of a picture book rewriting method for some embodiments of the present application;

[0100] Figure 14 A structure block diagram of a picture book rewriting apparatus for some embodiments of the present application;

[0101] Figure 15 A structure block diagram of a picture book continuation apparatus for some embodiments of the present application. DETAILED DESCRIPTION

[0102] The embodiments will be described in detail below with reference to the drawings. When the description below refers to attachments, assemblies, devices, elements, components, or the like, it should be understood that when there are two or more such attachments, assemblies, devices, elements, components or the like, the two or more can be either a unitary structure or separate structures. The embodiments described below are merely examples for implementing the present application and should not be interpreted in a manner limiting the scope of the present application. The present application can be implemented in various manners, and the embodiments described below are merely examples.

[0103] It should be noted that the brief description of the terms in the present application is only for the convenience of understanding the implementation described next, and is not intended to limit the implementation of the present application. Unless otherwise specified, these terms should be understood according to their ordinary and general meanings.

[0104] The terms "first", "second", "third", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar or similar objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that the terms used in this way can be interchanged under appropriate circumstances.

[0105] The terms "comprises", "comprising", "includes", "including", "has", "having" and their conjugates mean, when used in this document, that the mentioned features are included, but not to the exclusion of other features. In other words, the terms "comprises", "comprising", "includes", "including", "has", "having" and their conjugates mean that the mentioned features are included, but that other features are not excluded.

[0106] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software codes that can perform the function related to the component.

[0107] In the embodiments of the present application, the display device 200 refers to a device with the ability of picture display and data processing. For example, the display device 200 includes, but is not limited to, a smart television, a mobile terminal, a computer, a monitor, an advertising screen, a wearable device, a virtual reality device, an augmented reality device, etc.

[0108] Figure 1 The schematic diagram of the operation scenario between the display device and the control device is provided for some embodiments of the present application. As shown in Figure 1 , the user can operate the display device 200 through a touch operation, a mobile terminal 300 and a control device 100. For example, the control device 100 can be a remote controller, a stylus, a handle, etc.

[0109] The mobile terminal 300 can be used as a kind of control device to perform the human-computer interaction between the user and the display device 200. The mobile terminal 300 can also be used as a kind of communication device to establish a communication connection with the display device 200 and perform data interaction. In some embodiments, the mobile terminal 300 can install a software application with the display device 200, realize the connection communication through a network communication protocol, and achieve the purpose of one-to-one control operation and data communication. The mobile terminal 300 can also display audio and video content on the display device 200 to realize the function of synchronous display.

[0110] As shown in Figure 1 , the display device 200 also communicates data with the server 400 through various communication modes. The display device 200 can be allowed to communicate through a local area network (LAN), a wireless local area network (WLAN) and other networks.

[0111] The display device 200 can provide a broadcast receiving television function, and can additionally provide a smart network television function with computer support function, including but not limited to a network television, a smart television, an Internet protocol television (IPTV), etc.

[0112] Figure 2 The hardware configuration block diagram of the display device 200 is shown in Figure 1

[0113] ​In some embodiments, the display device 200 can include at least one of a tuner and demodulator 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, a user input interface.

[0114] In some embodiments, the detector 230 is configured to collect signals of an external environment or an external interaction. For example, the detector 230 includes a light receiver configured to collect an ambient light intensity, or the detector 230 includes an image collector such as a camera configured to collect an external environment scene, a user attribute, or a user interaction gesture, or the detector 230 includes a sound collector such as a microphone configured to receive an external sound.

[0115] In some embodiments, the display 260 includes a display functional component configured to present a picture, and a driving component configured to drive an image display. The display 260 is configured to receive an image signal output from the controller 250 for display. For example, the display 260 can be configured to display a video content, an image content, and a component of a menu control interface, and a user control UI interface.

[0116] In some embodiments, the communication device 220 is a component configured to communicate with an external device or a server 400 according to various communication protocol types. The display device 200 can be provided with multiple communication devices 220 according to different supported communication manners. For example, when the display device 200 supports wireless network communication, the display device 200 can be provided with a communication device 220 including a WiFi function. When the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 including a Bluetooth function.

[0117] The communication device 220 can be configured to connect the display device 200 to the external device or the server 400 in a wireless or wired manner. The wired connection can be achieved by connecting the display device 200 to the external device through a data line, an interface, or the like. The wireless connection can be achieved by connecting the display device 200 to the external device through a wireless signal or a wireless network. The display device 200 can be directly connected to the external device, or can be indirectly connected to the external device through a gateway, a router, a connection device, or the like.

[0118] In some embodiments, the controller 250 can include at least one of a central processor, a video processor, an audio processor, a graphics processor, a power supply processor, a first interface to an n-th interface for input / output, and the controller 250 can control the operation of the display device and respond to the user's operation by controlling various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200.

[0119] In some embodiments, the controller 250 and the tuner demodulator 210 can be located in different split devices, i.e. the tuner demodulator 210 can also be located in an external device of the main device where the controller 250 is located, such as an external set-top box, etc.

[0120] In some embodiments, the user can input user commands through a graphical user interface (GUI) displayed on the display 260, and the user input interface receives the user input commands through the graphical user interface (GUI).

[0121] In some embodiments, the audio output device 270 can be a native loudspeaker of the display device 200, or can be an audio output device connected to the display device 200. For the audio output device connected to the display device 200, the display device 200 can also be provided with an external audio output terminal, and the audio output device can be connected to the display device 200 through the external audio output terminal to output the sound of the display device 200.

[0122] In some embodiments, the user input interface 280 can be used to receive instructions from the user input.

[0123] Figure 3 The hardware configuration block diagram of the control device provided in some embodiments of the present application is shown in Figure 1 FIG. 1. As shown in Figure 3 The control device 100 can include a controller 110, a communication interface 130, a user input / output interface, a memory, and a power supply.

[0124] The control device 100 is configured to control the display device 200, and can receive the input operation instructions of the user, and convert the operation instructions into instructions that the display device 200 can recognize and respond to, thereby playing the role of an intermediary between the user and the display device 200.

[0125] In some embodiments, the control device 100 can be a smart device. For example, the control device 100 can install various applications for controlling the display device 200 according to the user's needs.

[0126] In some embodiments, as shown in Figure 1 the mobile terminal 300 or other smart electronic devices can play a similar function to the control device 100 after installing the application for controlling the display device 200.

[0127] The controller 110 includes a processor 112 and a RAM 113 and a ROM 114, a communication interface 130, and a communication bus. The controller 110 is used to control the operation and operation of the control device 100, and the communication and cooperation between the internal components, as well as the data processing function of the external and internal.

[0128] The communication interface 130, under the control of the controller 110, implements communication of control signals and data signals with the display device 200. The communication interface 130 can include at least one of a WiFi chip 131, a Bluetooth module 132, an NFC module 133, and other near field communication modules.

[0129] The user input / output interface 140, wherein the input interface includes at least one of a microphone 141, a touchpad 142, a sensor 143, a key 144, and other input interfaces.

[0130] In some embodiments, the control device 100 includes at least one of the communication interface 130 and the input / output interface 140. The control device 100 is configured with a communication interface 130, such as a WiFi, Bluetooth, NFC, etc. module, which can encode user input instructions through a WiFi protocol, or a Bluetooth protocol, or an NFC protocol and send them to the display device 200.

[0131] The storage 190 is used to store various running programs, data and applications for driving and controlling the control device 100 under the control of the controller. The storage 190 can store various control signal instructions input by the user.

[0132] The power supply 180 is used to provide operating power support for the elements of the control device 100 under the control of the controller.

[0133] In order to perform user interaction, in some embodiments, the display device 200 can run an operating system. The operating system is a computer program for managing and controlling hardware resources and software resources in the display device 200. The operating system can provide a user interface to allow the user to interact with the display device 200 and support the running of various application programs.

[0134] It should be noted that the operating system can be a native operating system based on a specific operating platform, a third-party operating system deeply customized based on a specific operating platform, or an independent operating system specially developed for the display device.

[0135] The operating system can be divided into different modules or levels according to the functions implemented, for example, as shown in FIG. 1C, in some embodiments, the system is divided into four layers, from top to bottom, the application layer (referred to as "application layer"), the application framework layer (referred to as "framework layer"), the system library layer and the kernel layer. Figure 4

[0136] ​In some embodiments, the application layer provides services and interfaces for applications, enabling the display device 200 to run applications and interact with the user based on the applications. The application layer may contain at least one application, which may be a built-in Windows program, system settings program, or clock program of the operating system; or it may be an application developed by a third-party developer. In specific implementations, the application packages in the application layer are not limited to the examples above.

[0137] The framework layer provides application programming interfaces (APIs) and a programming framework for applications. The application framework layer includes predefined functions. It acts as a central processing unit, determining the actions taken by applications within the application layer. Through the API, applications can access system resources and obtain system services during execution.

[0138] like Figure 4 As shown, the application framework layer in this embodiment includes a view system, managers, and content providers. The view system designs and implements the application's interface and interactions, and includes lists, grids, text boxes, and buttons. The managers include at least one of the following modules: an activity manager for interacting with all running activities in the system; a location manager for providing system services or applications with access to system location services; a package manager for retrieving various information related to application packages currently installed on the device; a notification manager for controlling the display and clearing of notification messages; and a window manager for managing icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.

[0139] In some embodiments, the Activity Manager manages the lifecycle of individual applications and common navigation and back functions, such as controlling application exit, opening, and back actions. The Window Manager manages all window programs, such as obtaining the screen size, determining if a status bar is present, locking the screen, capturing the screen, and controlling changes to the display window, such as shrinking the display window, shaking the display, or distorting the display.

[0140] In some embodiments, the system runtime layer can provide support for the framework layer. When the framework layer is used, the operating system runs the instruction library contained in the system runtime layer, such as the C / C++ instruction library, to implement the functions implemented by the framework layer.

[0141] In some embodiments, the kernel layer is a functional layer between the hardware and the software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, memory management, etc. For example, as shown in Figure 4 The kernel layer can be configured with hardware drivers. The drivers contained in the kernel layer can be at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power supply driver, etc.

[0142] It should be noted that the above examples are only a simple division of the functions of the operating system, and do not constitute a limitation on the specific operating system form of the display device 200 in the embodiments of the present application. According to the function of the display device, the type of the operating system, and other factors, the number and specific type of the layers contained in the operating system can be in other forms.

[0143] In some embodiments, as shown in Figure 5 The present application provides a picture book rewriting method applied to the display device 200. The method can include the following steps:

[0144] Step S501, in the case of displaying picture book content on the user interface, receiving rewriting requirement information for the picture book content, and taking the identifier of the current picture book as the picture book identifier corresponding to the rewriting requirement information;

[0145] Step S502, identifying the text information corresponding to the rewriting requirement information, and obtaining the picture book information corresponding to the picture book identifier from the database according to the picture book identifier corresponding to the rewriting requirement information; the picture book information represents the multi-dimensional information of the picture book corresponding to the picture book identifier;

[0146] Step S503, inputting the text information and the information representing the role in the picture book information into the text processing model to obtain the rewritten text content; the rewritten text content represents the text content composed of the text information and the information representing the role; the text processing model is used to perform picture book text content generation processing according to the input text information and the information representing the role, and output the rewritten text content;

[0147] In step S504, the rewritten text content, the information representing the style in the picture book information, and the information representing the character features in the picture book information are input into an image processing model to obtain rewritten image data. The rewritten image data represents image data matched with the rewritten text content, the information representing the style, and the information representing the character features. The image processing model is used to perform picture book image data generation processing according to the input rewritten text content, the information representing the style, and the information representing the character features, and output the rewritten image data.

[0148] In step S505, based on the rewritten text content and the information representing the voice in the picture book information, a play audio data synthesis processing is performed to obtain play audio data corresponding to the rewritten image data, and based on the information representing the theme and the information representing the style in the picture book information, a background audio data generation processing is performed to obtain background audio data corresponding to the rewritten image data.

[0149] In step S506, the picture book image data in the picture book content in the picture book corresponding to the picture book identifier is replaced with the rewritten image data, and the play audio data and the background audio data corresponding to the picture book image data are replaced with the play audio data and the background audio data corresponding to the rewritten image data to obtain a rewritten picture book. The number of pages of the rewritten picture book and the number of pages of the picture book corresponding to the picture book identifier remain unchanged.

[0150] The user interface can be an interactive interface for displaying picture book content and receiving user operations, for example, a display interface of a picture book.

[0151] The picture book content can be a combination of image and text content, for example, a combination of story plot and text content.

[0152] The rewriting requirement information can be a modification request made by the user for the picture book content, for example, a modification content input by the user through voice, such as “rewrite this story content as the little friend crying loudly and knocking the medicine bottle to the ground”.

[0153] The picture book identifier can be identification information for uniquely identifying a picture book, for example, a unique identifier generated for each story.

[0154] The database can be a data storage system for storing picture book related information, for example, a database system for storing character information, style information, story theme, story content, voice ID, and character reference features.

[0155] The text information can be specific text content identified from the rewriting requirement information, for example, text content converted from user voice input.

[0156] The picture book information can be all information related to the picture book stored in the database, for example, can be character information, style information, story theme, story content, voice color ID, character reference features, and the like.

[0157] The multi-dimensional information can be a set of information describing the characteristics of various aspects of the picture book, for example, can be information containing multiple dimensions of characters, styles, themes, content, and voice colors.

[0158] The text processing model can be a model used for text generation tasks, for example, can be a large language model.

[0159] The rewritten text content can be new text content generated by the model according to user input and character information, for example, can be new story content that meets the user's rewriting requirements and maintains the consistency of the character.

[0160] The information representing the character can be information describing the characteristics of the story character, for example, can be character name, character attribute, and the like.

[0161] The image processing model can be a model used to generate images, for example, can be a text-to-image large model.

[0162] The rewritten image data can be a new image generated according to the rewritten text content, for example, can be an image generated by a text-to-image model that matches the new story content.

[0163] The information representing the style can be information describing the style of the picture book, for example, can be cartoon, realistic, hand-drawn, watercolor, and the like.

[0164] The information representing the character features can be feature information used to maintain the consistency of the character, for example, can be character attribute-related features extracted from the text-to-image model level.

[0165] The information representing the voice can be voice information used for speech synthesis, for example, can be system-integrated voice colors such as little girl, little boy, mature male, mature female, and the like.

[0166] The voice audio data can be voice data converted from the text content, for example, can be story voice audio generated by a speech synthesis model.

[0167] The information representing the theme can be information describing the theme of the story, for example, can be overall theme style information of the story.

[0168] The background audio data can be background music data matched with the story, for example, can be light music generated according to the story theme and style.

[0169] Specifically, the controller 250 receives the rewriting requirement information input by the user through the voice recognition module when the user interface displays the picture book content, and generates a unique picture book identifier for the current picture book; the controller 250 inputs the rewriting requirement information into the large language model for voice-to-text processing to obtain text information, and simultaneously obtains picture book information corresponding to the picture book identifier from the database, the picture book information including role information, style information, theme information, and broadcast timbre information; the controller 250 inputs the text information and the role-representing information in the picture book information into the large language model to generate rewritten text content conforming to the role setting; the controller 250 inputs the rewritten text content, the style-representing information, and the role feature-representing information in the picture book information into the text-to-image large model, and generates rewritten image data by maintaining the consistency of the role image through the consistency attention mechanism; the controller 250 synthesizes broadcast audio data based on the rewritten text content and the broadcast timbre-representing information using the multi-modal generation large model, and generates background audio data based on the theme-representing information and the style-representing information; finally, the controller 250 replaces the images, the broadcast audio, and the background audio in the original picture book content with the newly generated content to obtain a rewritten picture book, while keeping the number of pages unchanged.

[0170] The technical scheme provided in this embodiment is advantageous in that the picture book identifier is associated with multi-dimensional information and stored in the database, and the role, style, theme, and the like are uniformly obtained based on the picture book identifier during rewriting, which helps to avoid redundant operations such as repeated acquisition and manual input of information; the obtained information is input into a text processing model, an image processing model, and an audio processing model for automatic processing, and accurate replacement is performed while keeping the number of pages unchanged, which helps to maintain the consistency of the rewritten content and the original picture book in terms of role, style, theme, and the like; thereby, the automatic rewriting of the picture book is facilitated, the manual operation steps are reduced, and the rewriting efficiency of the picture book is improved.

[0171] In some embodiments, as shown in Figure 6 In some embodiments, as shown in

[0172] Step S601, in the case where the user interface displays picture book content, receiving a picture book content selection operation input by the user.

[0173] Step S602, in the case where the selected picture book content is identified according to the picture book content selection operation, receiving rewriting requirement information for the picture book content.

[0174] The picture book content selection operation can be an operation of selecting the picture book content by the user through an input device, for example, an operation of focusing on a certain picture book image by the user through a remote controller.

[0175] For example, referring to Figure 7The display 260 displays a picture book image and picture book text content of a picture book story page of a picture book story in a user interface of the picture book application. The picture book image can refer to Figure 7 a child holding a medicine in Figure 7 , the picture book text content can refer to a text part in Figure 7 , such as "The child smiles, he stretches out his hand, gently takes the medicine, and prepares to eat it", and the picture book story page can be a page 8 picture book story page in a picture book story with a theme of "The Adventure Journey of the Medicine". The user can also return to the home page by clicking the back key.

[0176] Specifically, the controller 250 receives a picture book content selection operation input by the user through the remote controller when displaying the picture book content in the user interface. When the user moves the focus to a certain page of the picture book content, the controller 250 identifies the specific picture book content selected by the user and activates the microphone of the voice receiving module. After the controller 250 identifies the specific picture book content selected by the user according to the picture book content selection operation, the controller 250 receives the rewriting requirement information input by the user for the selected picture book content through the voice receiving module.

[0177] The technical solution provided in this embodiment receives the selection operation of the user on the displayed picture book content first, identifies the specific picture book content based on the selection operation, and then receives the rewriting requirement information for the selected content. This is beneficial to accurately positioning the specific content position that the user wants to rewrite and avoiding the problem of unclear rewriting range. Therefore, it is beneficial to ensure that there is a clear corresponding relationship between the rewriting requirement information and the picture book content selected by the user and improve the accuracy of picture book rewriting.

[0178] In some embodiments, as shown in Figure 8 , after receiving the input picture book content selection operation, the method further includes:

[0179] Step S801: In response to the picture book content selection operation, identifying the picture book content corresponding to the current focus;

[0180] Step S802: Identifying the picture book content corresponding to the current focus as the selected picture book content.

[0181] The current focus can be the position currently selected by the remote controller on the user interface.

[0182] The picture book content corresponding to the current focus can be the picture book image displayed at the position of the current focus of the remote controller, for example, it can be the picture book image displayed on the page 8 of the picture book story when the user places the focus on the picture book image displayed on the page 8 of the picture book story through the remote controller.

[0183] Specifically, the controller 250 receives the picture book content selection operation input by the user through the remote controller through the application control module, and monitors the focus position of the remote controller on the user interface of the picture book application in real time. When detecting that the focus position of the remote controller changes, the controller 250 obtains coordinate information of the current focus position, and locates the picture book content corresponding to the coordinate information in the user interface of the picture book application according to the coordinate information. After the controller 250 locates the picture book content corresponding to the current focus, the controller 250 takes the picture book content as the picture book content selected by the user, so as to subsequently receive the rewriting requirement information for the picture book content.

[0184] The technical scheme provided by the embodiment is advantageous in accurately positioning the specific content position that the user wants to rewrite, and avoiding misidentifying the rewriting intention of the user, by responding to the picture book content selection operation and identifying the picture book content corresponding to the current focus. The technical scheme is advantageous in establishing an accurate correspondence between the focus position and the specific content, by identifying the picture book content corresponding to the current focus as the selected picture book content. Therefore, the technical scheme is advantageous in accurately understanding the rewriting requirement of the user, and improving the accuracy of picture book rewriting.

[0185] In some embodiments, before obtaining the picture book information corresponding to the picture book identifier of the rewriting requirement information from the database, the method further includes: performing intent recognition processing on the text information corresponding to the rewriting requirement information to obtain an intent recognition result corresponding to the rewriting requirement information; the intent recognition result is used to indicate whether to obtain data from the database; determining whether to obtain data from the database based on the intent recognition result; and obtaining the picture book information corresponding to the picture book identifier of the rewriting requirement information from the database, including: in a case where it is determined to obtain data from the database, obtaining the picture book information corresponding to the picture book identifier of the rewriting requirement information from the database.

[0186] The intent recognition processing can be a processing process of analyzing and understanding the rewriting requirement information input by the user.

[0187] The intent recognition result can be a classification result obtained after the intent recognition processing on the rewriting requirement information input by the user, for example, a result of judging that the user wants to rewrite the picture book content and needs to obtain related data from the database.

[0188] Specifically, after receiving the rewriting demand information, the controller 250 sends text information corresponding to the rewriting demand information to the domain intention recognition module for intention recognition processing; the controller 250 receives an intention recognition result returned by the domain intention recognition module, and the intention recognition result is used to represent whether data needs to be obtained from a database; the controller 250 determines whether data needs to be obtained from the database based on the intention recognition result, and when it is determined that data needs to be obtained from the database, the controller 250 obtains, from the database, book information such as character information, style information, story theme, story content, voice ID, and character reference features corresponding to a book identification according to the rewriting demand information.

[0189] The technical scheme provided in this embodiment is advantageous in avoiding unnecessary database access operations by performing intention recognition processing on text information corresponding to rewriting demand information and determining whether data needs to be obtained from a database based on an intention recognition result, and is advantageous in realizing on-demand execution of data obtaining operations by obtaining book information from the database only when it is really necessary, so as to facilitate reduction of resource consumption and improvement of overall operation efficiency while ensuring normal implementation of the rewriting function.

[0190] In some embodiments, the information representing the character features is also obtained in the following manner: book image data containing information representing the character is selected from a book corresponding to the book identification as reference book image data; and information representing the character features is extracted from the reference book image data.

[0191] The book image data can be a generated image frame containing a specific character in the book, for example, can be image data corresponding to a page of the book containing the “Yao Yao” character.

[0192] The reference book image data can be an image frame selected from the book for extracting the character features, for example, can be an image frame selected from a plurality of image frames containing the “Yao Yao” character, used to maintain consistency of the character features in the subsequent generation process.

[0193] Specifically, the controller 250 selects a book image data containing information representing the character as reference book image data from a book corresponding to the book identification by using a random sampling function; the controller 250 inputs the reference book image data into a U-Net diffusion model, extracts tokens features through an attention layer; and the controller 250 stores the extracted tokens features as information representing the character features in the database, used to maintain consistency of the character features in the subsequent generation process.

[0194] The technical scheme provided by the embodiment is advantageous in obtaining the standardized visual features of the character by selecting the picture book image data containing the character information as the reference from the picture book, and advantageous in establishing the reference data of the character features by extracting the information representing the character features from the reference picture book image data, thereby being advantageous in maintaining the consistency of the character image and improving the picture book generation quality in the subsequent picture book generation process.

[0195] In some embodiments, the information representing the character in the text information and the picture book information is input into the text processing model to obtain rewritten text content, including: inputting the text information into the text processing model to obtain key text information; the key text information represents the key information in the text information; the text processing model is used to output the key text information in the text information according to the input text information; inputting the key text information and the information representing the character into the text processing model to obtain the rewritten text content; the rewritten text content represents the text content composed of the key text information and the information representing the character; the text processing model is used to perform picture book text content generation processing according to the input key text information and the information representing the character, and output the rewritten text content.

[0196] The key text information can be the core information content in the user input text after being processed by the text processing model, for example, can be the key information such as “children crying” and “medicine bottle overturned” extracted from the content “Rewrite this story content as: Children cry and overturn the medicine bottle on the ground”.

[0197] Specifically, the controller 250 inputs the user input text information into the large language model for processing, extracts the core content in the text information through the key information extraction module to obtain the key text information; the controller 250 inputs the extracted key text information and the information representing the character obtained from the database into the large language model, and the large language model generates new picture book text content according to the input information, and converts the generated picture book text content into content suitable for image generation, thereby obtaining the rewritten text content.

[0198] The technical scheme provided by the embodiment is advantageous in accurately grasping the core content of the user's rewriting requirement by inputting the text information into the text processing model to extract the key text information first, and advantageous in generating new text content on the basis of maintaining the character features by inputting the extracted key text information and the information representing the character into the text processing model, thereby being advantageous in maintaining the coherence of the picture book character image while ensuring that the rewritten content meets the user's requirements, and improving the generation quality of the rewritten text content.

[0199] In some embodiments, inputting the rewritten text content, the information representing the style in the picture book information, and the information representing the character features in the picture book information into the image processing model to obtain the rewritten image data comprises: inputting the rewritten text content into a text processing model to obtain image generation prompt information corresponding to the rewritten text content; the image generation prompt information is used to represent text content adapted to the generation processing of the picture book image data; the text processing model is used to adjust the input rewritten text content and output the adjusted text content; the adjusted text content is the image generation prompt information; inputting the image generation prompt information, the information representing the style, and the information representing the character features into the image processing model to obtain the rewritten image data; the rewritten image data represents image data matched with the image generation prompt information, the information representing the style, and the information representing the character features; the image processing model is used to perform the generation processing of the picture book image data according to the input image generation prompt information, the information representing the style, and the information representing the character features, and output the rewritten image data.

[0200] The image generation prompt information can be a text description that is more suitable for image generation after adjustment and optimization, for example, it can be a more specific and detailed descriptive text such as "the child is very happy" rewritten as "the child holds up his hands and his face is full of smiles".

[0201] Specifically, the controller 250 inputs the rewritten text content into a large language model for processing, and converts a simple description such as "the child is very happy" into a detailed prompt word more suitable for image generation; the controller 250 obtains the information representing the style and the information representing the character features from the database, and inputs the converted image generation prompt information, the information representing the style, and the information representing the character features into the text-to-image large model; the controller 250 processes the input information through a consistency attention module CAB in a U-Net diffusion model, and maintains the consistency of the character features in the down-sampling layer and the up-sampling layer by using the cross-attention mechanism, and finally outputs the rewritten image data.

[0202] The technical scheme provided in the embodiment is advantageous in improving the accuracy of text-to-image conversion by converting the rewritten text content into prompt information more suitable for image generation; it is advantageous to maintain the consistency of the style and the character features in the image generation process by inputting the image generation prompt information, the information representing the style, and the information representing the character features into the image processing model together; thereby it is advantageous to generate high-quality image data that meets the user's rewriting requirements and maintains the original picture book style and character features.

[0203] In some embodiments, the image generation prompt information, the information representing the style, and the information representing the character features are input into an image processing model to obtain rewritten image data, including: inputting the image generation prompt information and the information representing the character features into the image processing model to obtain initial rewritten image data; the initial rewritten image data represents image data matched with the image generation prompt information and the information representing the character features; the image processing model is used to perform picture book image data generation processing according to the input image generation prompt information and the information representing the character features, and output the initial rewritten image data; inputting the initial rewritten image data and the information representing the style into the image processing model to obtain the rewritten image data; the style information of the rewritten image data is the same as the information representing the style, and the image content of the rewritten image data is the same as the image content of the initial rewritten image data; the image processing model is used to perform rendering processing of the information representing the style on the input initial rewritten image data, and output the rewritten image data.

[0204] In some embodiments, the initial rewritten image data can be original image data generated by the image processing model according to the image generation prompt information and the information representing the character features, which has not been subjected to style rendering.

[0205] Specifically, the controller 250 first inputs the image generation prompt information and the information representing the character features into the text-to-image large model, and processes in the down-sampling layer and the up-sampling layer of the U-Net diffusion model through the consistency attention module CAB; the controller 250 selects reference image features containing the same character from the database through the RandSample function, and calculates the cross attention between the current generated image features and the reference image features in the U-Net diffusion model to generate initial rewritten image data; the controller 250 inputs the generated initial rewritten image data and the information representing the style into the text-to-image large model together, and performs style rendering processing on the initial rewritten image data to obtain rewritten image data with a specified style.

[0206] The technical scheme provided by the embodiment divides the generation process of the rewritten image data into two steps, first generates the initial rewritten image data according to the image generation prompt information and the information representing the character features, which is beneficial to ensure the content accuracy of the generated image and the consistency of the character features; and then performs style rendering processing on the initial rewritten image data to obtain the rewritten image data, which is beneficial to realize the unity of the style while keeping the image content unchanged; thereby, it is beneficial to improve the generation quality of the rewritten image data, so that the generated image can accurately express the content and maintain the consistency of the style.

[0207] In some embodiments, based on the rewritten text content and the information representing the tone of the broadcast in the picture book information, the playback audio data synthesis processing is performed to obtain the broadcast audio data corresponding to the rewritten image data, and based on the information representing the theme and the information representing the style in the picture book information, the background audio data generation processing is performed to obtain the background audio data corresponding to the rewritten image data, including: inputting the rewritten text content and the information representing the tone of the broadcast into a playback audio data processing model to obtain the broadcast audio data corresponding to the rewritten image data; the tone information of the broadcast audio data is the same as the information representing the tone of the broadcast, and the text content corresponding to the broadcast audio data is the same as the rewritten text content; the playback audio data processing model is used to convert the input rewritten text content into corresponding initial broadcast audio data, adjust the initial broadcast audio data according to the information representing the tone of the broadcast, and output the broadcast audio data corresponding to the rewritten image data; inputting the information representing the theme and the information representing the style into a background audio data processing model to obtain the background audio data corresponding to the rewritten image data; the background audio label of the background audio data is the background audio label corresponding to the information representing the theme and the information representing the style; the background audio data processing model is used to generate background audio data according to the background audio label corresponding to the input information representing the theme and the information representing the style, and output the background audio data corresponding to the rewritten image data.

[0208] Among them, the playback audio data processing model can be an artificial intelligence model for converting text content into speech, for example, it can be a multi-modal generation large model.

[0209] Among them, the initial broadcast audio data can be voice data that has not been adjusted in tone after the text is converted into speech by the playback audio data processing model, for example, it can be voice data converted from text content into a default tone.

[0210] Among them, the background audio data processing model can be an artificial intelligence model for generating background music, for example, it can be a multi-modal generation large model.

[0211] Among them, the background audio label can be identification information for specifying the type of background music, for example, it can be a music style label such as cheerful, relaxed, etc.

[0212] Specifically, the controller 250 inputs the rewritten text content and the information representing the tone of the broadcast into the multi-modal generation large model, first converts the rewritten text content into initial broadcast audio data, and then performs tone adjustment processing on the initial broadcast audio data according to the information representing the tone of the broadcast, to generate broadcast audio data corresponding to the rewritten image data; the controller 250 simultaneously inputs the information representing the theme and the information representing the style into the multi-modal generation large model, determines the corresponding background audio label according to the story theme and the style feature, generates background audio data matched with the rewritten image data through a background audio data processing model, and finally performs audio mixing processing on the broadcast audio data and the background audio data to form complete audio output.

[0213] The technical scheme provided by the embodiment separates the generation processes of the broadcast audio data and the background audio data, which is conducive to controlling the tone feature of the broadcast content and the style feature of the background music respectively; by performing tone adjustment processing on the broadcast audio data and theme and style label matching processing on the background audio data, the consistency of the generated audio with the rewritten content, theme and style can be ensured.

[0214] In some embodiments, with reference to Figure 9 The present application also provides a picture book continuation method, which is applied to the display device 200, and can include the following steps:

[0215] In step S901, in the case that the picture book is displayed on the user interface, the continuation demand information for the picture book is received, and the identifier of the current picture book is taken as the picture book identifier corresponding to the continuation demand information.

[0216] In step S902, the text information corresponding to the continuation demand information is identified, and the picture book information corresponding to the picture book identifier is obtained from the database according to the picture book identifier corresponding to the continuation demand information; the picture book information represents the multi-dimensional information of the picture book corresponding to the picture book identifier.

[0217] In step S903, the information representing the character in the text information and the picture book information is input into a text processing model to obtain the continuation text content; the continuation text content represents the text content composed of the text information and the information representing the character; the text processing model is used to perform picture book text content generation processing according to the input text information and the information representing the character, and output the continuation text content.

[0218] In step S904, the continuation text content, the information representing the style in the picture book information, and the information representing the character features in the picture book information are input into an image processing model to obtain continuation image data. The continuation image data represents image data matched with the continuation text content, the information representing the style, and the information representing the character features. The image processing model is used to perform picture book image data generation processing according to the input continuation text content, the information representing the style, and the information representing the character features, and output the continuation image data.

[0219] In step S905, based on the continuation text content and the information representing the voice in the picture book information, a play audio data synthesis processing is performed to obtain play audio data corresponding to the continuation image data, and based on the information representing the theme and the information representing the style in the picture book information, a background audio data generation processing is performed to obtain background audio data corresponding to the continuation image data.

[0220] In step S906, semantic recognition processing is performed on the text information corresponding to the continuation requirement information to obtain a continuation position of the continuation image data in a picture book corresponding to the picture book identifier. The continuation image data is added in the continuation position in the picture book corresponding to the picture book identifier, and the play audio data and the background audio data corresponding to the continuation image data are configured to obtain a continuation picture book. The number of pages of the continuation picture book is greater than the number of pages of the picture book corresponding to the picture book identifier.

[0221] The continuation requirement information can be a content expansion request information proposed by a user for an existing picture book, for example, the user can input a continuation request such as “At the end of the story, the pill invented a magical medicine, and after eating it, the children will never get sick” through voice input.

[0222] The continuation text content can be new story content generated by a text processing model based on user input text information and original character information, for example, a new plot description can be generated based on maintaining the original character features.

[0223] The continuation image data can be new image data generated by an image processing model based on the continuation text content, the style information, and the character feature information, for example, it can be a picture book image that maintains the original character image and artistic style corresponding to the new plot.

[0224] The continuation position can be an insertion position of the new content in the original picture book determined by semantic recognition, for example, the continuation position can be at the beginning of the picture book or at the end of the picture book.

[0225] The continuation picture book can be a new picture book formed by adding the continuation content to the original picture book, for example, it can be an 11-page picture book formed by adding 1 page of new content to the end of the original 10-page picture book.

[0226] Specifically, the controller 250 receives the user's continuation demand information when the user interface displays the picture book, and assigns a picture book identifier to the continuation demand information; the controller 250 identifies the text information in the continuation demand information through the large language model, and obtains the picture book information containing the character information, style information and theme information from the database; the controller 250 inputs the text information and the information representing the character into the large language model to generate the continuation text content, and then inputs the continuation text content, the information representing the style and the information representing the character features into the image-from-text large model, processes it through the consistency attention module CAB (consistency attention module) in the down-sampling layer and the up-sampling layer of the U-Net (U network) diffusion model, uses a random sampling function to select reference image features containing the same character from the database, and generates continuation image data by calculating cross attention; the controller 250 inputs the continuation text content and the information representing the broadcast tone into the multi-modal generation large model to generate broadcast audio data, and simultaneously generates background audio data according to the information representing the theme and the information representing the style; finally, the controller 250 determines the insertion position of the continuation content through semantic recognition, adds the continuation image data, the broadcast audio data and the background audio data to the original picture book, and forms a picture book with more pages after the continuation. For example, the contents of steps S902, S903, S904 and S905 can refer to the contents of steps S502, S503, S504 and S505 in the picture book rewriting method described above.

[0227] For example, referring to Figure 10 The display 260 displays a plurality of picture book stories in the user interface of the picture book application, which can include "The Wonderful Journey of Pill Pill", "The Wonderful Life of Kitten", "The Lovely World of Someone", etc. In this interface, the user can see the introduction of the picture book story, can click on the AI picture book generation to automatically generate the picture book, and can also delete the customized picture book by pressing the menu key.

[0228] The technical solution provided by the embodiment acquires multi-dimensional information of the original picture book and uses it as a constraint condition for continuation, which is conducive to ensuring that the continuation content is consistent with the original picture book in terms of character image, artistic style and story theme; by cooperatively generating text content, image data, broadcast audio and background audio, and determining the appropriate continuation position according to semantic recognition, it is conducive to realizing the natural connection of multi-modal content, expanding the content of the picture book while maintaining the coherence and integrity of the story; thereby facilitating the automatic continuation of the picture book, reducing manual operation steps, and thereby improving the continuation efficiency of the picture book.

[0229] In some embodiments, in the case that the user interface displays the picture book, receiving the writing demand information for the picture book comprises: in the case that the user interface displays the picture book, receiving an input picture book selection operation; in the case that a selected picture book is identified according to the picture book selection operation, receiving the writing demand information for the picture book.

[0230] The picture book selection operation can be an operation of selecting a specific picture book by the user through an input device, for example, can be an operation of controlling the focus to fall on a certain picture book by the user through a remote controller.

[0231] Specifically, the controller 250 displays the picture book list on the user interface, when receiving the picture book selection operation input by the user through the remote controller, the controller 250 positions the focus on the picture book selected by the user, and activates the voice input function; the controller 250 calls the voice receiving module to open the sound receiving microphone, and receives the writing demand information for the picture book.

[0232] The technical scheme provided by the embodiment, by displaying the picture book on the user interface and receiving the picture book selection operation of the user, is conducive to accurately positioning the specific picture book that needs to be written; by receiving the writing demand information after identifying the picture book selected by the user, is conducive to ensuring the accurate correspondence between the writing demand and the specific picture book; thereby, it is conducive to establishing an orderly picture book writing operation process, avoiding the wrong matching between the writing demand and the picture book, and improving the accuracy of the picture book writing.

[0233] In some embodiments, after receiving the input picture book selection operation, further comprising: in response to the picture book selection operation, identifying the picture book corresponding to the current focus; identifying the picture book corresponding to the current focus as the selected picture book.

[0234] The selected picture book can be a specific picture book selected by the user through the current focus position, for example, when the current focus falls on the picture book "Medicine Medicine's Adventure Journey", it is identified that the picture book "Medicine Medicine's Adventure Journey" is selected by the user.

[0235] Specifically, the controller 250 displays the picture book list on the user interface, when receiving the picture book selection operation input by the user through the remote controller, the controller 250 responds to the picture book selection operation, identifies the specific position of the current focus in the user interface through the focus positioning module, and identifies the picture book corresponding to the current focus from the displayed picture book list according to the position information of the current focus, and identifies it as the selected picture book.

[0236] For example, in the case of displaying a picture book on the user interface, receiving the continuation demand information for the picture book includes the following: in the case of displaying a picture book on the user interface, the controller 250 receives an input picture book selection operation; in response to the picture book selection operation, identifying the picture book corresponding to the current focus; identifying the picture book corresponding to the current focus as the selected picture book; in the case of identifying the selected picture book according to the picture book selection operation, receiving the continuation demand information for the picture book. By responding to the picture book selection operation of the user and identifying the picture book corresponding to the current focus position, it is beneficial to establish a clear correspondence between the user operation and the specific picture book; by identifying the picture book corresponding to the current focus as the picture book selected by the user, it is beneficial to accurately locate the specific picture book that the user wants to continue; thereby it is beneficial to improve the accuracy of the user's picture book selection operation; by receiving the continuation demand information after identifying the picture book selected by the user, it is beneficial to ensure the accurate correspondence between the continuation demand and the specific picture book; thereby it is beneficial to establish an orderly picture book continuation operation process, avoid the wrong matching between the continuation demand and the picture book, and improve the accuracy of the picture book continuation.

[0237] The technical scheme provided by the embodiment establishes a clear correspondence between the user operation and the specific picture book by responding to the picture book selection operation of the user and identifying the picture book corresponding to the current focus position; and accurately locates the specific picture book that the user wants to continue by identifying the picture book corresponding to the current focus as the picture book selected by the user; thereby it is beneficial to improve the accuracy of the user's picture book selection operation.

[0238] In some embodiments, in the picture book identified corresponding continuation position in the picture book, adding the continuation image data, and configuring the broadcast audio data and the background audio data corresponding to the continuation image data, obtaining the continued picture book, including: in the case that the continuation position is the end of the picture book identified corresponding picture book, adding the continuation image data after the last page of the picture book identified corresponding picture book, and configuring the broadcast audio data and the background audio data corresponding to the continuation image data, obtaining the continued picture book; in the case that the continuation position is the beginning of the picture book identified corresponding picture book, adding the continuation image data before the first page of the picture book identified corresponding picture book, and configuring the broadcast audio data and the background audio data corresponding to the continuation image data, obtaining the continued picture book.

[0239] Specifically, the controller 250 obtains character information, style information and story theme of the picture book from the database according to the picture book identification, and determines a content adding manner according to the continuation position; when the continuation position is the end of the picture book, the controller 250 adds the continuation image data after the last page of the picture book; when the continuation position is the beginning of the picture book, the controller 250 adds the continuation image data before the first page of the picture book; when generating the continuation image data, the controller 250 maintains the consistency of the character image through a consistency attention mechanism, and generates the broadcast audio data by calling a speech synthesis module according to the story theme, and generates the background audio data by calling a music generation module according to the picture book style, and finally integrates the continuation image data, the broadcast audio data and the background audio data to form the picture book after the continuation.

[0240] The technical scheme provided by the embodiment is advantageous in that the continuation image data is flexibly added at the beginning or the end of the picture book, so that the user can select a suitable continuation position according to actual needs; and the continuation image data is configured with corresponding broadcast audio data and background audio data, so that the consistency of the continuation content and the original picture book in visual and auditory experience is maintained.

[0241] In some embodiments, as shown in Figure 1 and 2 The present application also provides a display device 200, which can include a display 260 and a controller 250; wherein the controller 250 is coupled with the display 260. As shown in Figure 5 The controller 250 is configured to:

[0242] Step S501, in the case that the picture book content is displayed on the user interface, receiving rewriting requirement information for the picture book content, taking the identification of the current picture book as the picture book identification corresponding to the rewriting requirement information;

[0243] Step S502, identifying the text information corresponding to the rewriting requirement information, and obtaining the picture book information corresponding to the picture book identification from the database according to the picture book identification corresponding to the rewriting requirement information; the picture book information represents the multi-dimensional information of the picture book corresponding to the picture book identification;

[0244] Step S503, inputting the text information and the information representing the character in the picture book information into a text processing model to obtain rewritten text content; the rewritten text content represents the text content composed of the text information and the information representing the character; the text processing model is used for generating and processing the picture book text content according to the input text information and the information representing the character, and outputting the rewritten text content;

[0245] In step S504, the rewritten text content, the information representing the style in the picture book information, and the information representing the character features in the picture book information are input into an image processing model to obtain rewritten image data; the rewritten image data represents image data matched with the rewritten text content, the information representing the style, and the information representing the character features; the image processing model is used to perform picture book image data generation processing according to the input rewritten text content, the information representing the style, and the information representing the character features, and output the rewritten image data.

[0246] In step S505, based on the rewritten text content and the information representing the voice in the picture book information, playing audio data synthesis processing is performed to obtain playing audio data corresponding to the rewritten image data, and based on the information representing the theme and the information representing the style in the picture book information, background audio data generation processing is performed to obtain background audio data corresponding to the rewritten image data.

[0247] In step S506, the picture book image data in the picture book content in the picture book corresponding to the picture book identifier is replaced with the rewritten image data, and the playing audio data and the background audio data corresponding to the picture book image data are replaced with the playing audio data and the background audio data corresponding to the rewritten image data to obtain a rewritten picture book; the number of pages of the rewritten picture book and the number of pages of the picture book corresponding to the picture book identifier remain unchanged.

[0248] It should be noted that the specific implementation process of the display device 200 described above can refer to the related embodiments of the picture book rewriting method shown in FIG. 13, and will not be described here again. Figure 5

[0249] The technical scheme provided in this embodiment is beneficial to avoiding redundant operations of repeated information acquisition and manual input, by associating and storing the picture book identifier and the multi-dimensional information in the database, and uniformly acquiring information such as characters, styles, and themes based on the picture book identifier during rewriting; it is beneficial to keeping the consistency of the rewritten content and the original picture book in terms of characters, styles, and themes, by inputting the acquired information into a text processing model, an image processing model, and an audio processing model for automatic processing, and accurately replacing under the condition of keeping the number of pages unchanged; thereby, it is beneficial to realizing automatic rewriting of the picture book, reducing manual operation steps, and thereby improving the rewriting efficiency of the picture book.

[0250] In some embodiments, the controller 250 of the display device 200 described above can further perform the following steps:

[0251] In the case of displaying the picture book content on the user interface, the input picture book content selection operation is received; in the case of identifying the selected picture book content according to the picture book content selection operation, the rewriting requirement information of the picture book content is received.

[0252] ​In some embodiments, the controller 250 of the display device 200 described above can further perform the following steps:

[0253] In response to the picture book content selection operation, the picture book content corresponding to the current focus is identified; the picture book content corresponding to the current focus is identified as the selected picture book content.

[0254] In some embodiments, the controller 250 of the display device 200 described above can further perform the following steps:

[0255] The text information corresponding to the rewriting requirement information is subjected to intent recognition processing to obtain an intent recognition result corresponding to the rewriting requirement information; the intent recognition result is used to represent whether to obtain data from the database; based on the intent recognition result, it is judged whether to obtain data from the database; in the case of obtaining data from the database, the picture book information corresponding to the picture book identifier is obtained from the database according to the picture book identifier corresponding to the rewriting requirement information.

[0256] In some embodiments, the controller 250 of the display device 200 described above can further perform the following steps:

[0257] From the picture book corresponding to the picture book identifier, picture book image data containing information representing the role is selected as reference picture book image data; information representing the role characteristics is extracted from the reference picture book image data.

[0258] In some embodiments, the controller 250 of the display device 200 described above can further perform the following steps:

[0259] The text information is input into a text processing model to obtain key text information; the key text information represents the key information in the text information; the text processing model is used to output the key text information in the text information according to the input text information; the key text information and the information representing the role are input into the text processing model to obtain the rewritten text content; the rewritten text content represents the text content composed of the key text information and the information representing the role; the text processing model is used to perform picture book text content generation processing according to the input key text information and the information representing the role, and output the rewritten text content.

[0260] In some embodiments, the controller 250 of the display device 200 described above can further perform the following steps:

[0261] The rewritten text content is input into a text processing model to obtain image generation prompt information corresponding to the rewritten text content; the image generation prompt information is used to represent text content adapted to drawing image data generation processing; the text processing model is used to adjust the input rewritten text content and output adjusted text content; the adjusted text content is the image generation prompt information; the image generation prompt information, the information representing the style, and the information representing the character features are input into an image processing model to obtain rewritten image data; the rewritten image data represents image data matched with the image generation prompt information, the information representing the style, and the information representing the character features; the image processing model is used to perform drawing image data generation processing according to the input image generation prompt information, the information representing the style, and the information representing the character features, and output the rewritten image data.

[0262] In some embodiments, the controller 250 of the display device 200 described above can further perform the following steps:

[0263] The image generation prompt information and the information representing the character features are input into an image processing model to obtain initial rewritten image data; the initial rewritten image data represents image data matched with the image generation prompt information and the information representing the character features; the image processing model is used to perform drawing image data generation processing according to the input image generation prompt information and the information representing the character features, and output the initial rewritten image data; the initial rewritten image data and the information representing the style are input into the image processing model to obtain the rewritten image data; the style information of the rewritten image data is the same as the information representing the style, and the image content of the rewritten image data is the same as the image content of the initial rewritten image data; the image processing model is used to perform rendering processing of the information representing the style on the input initial rewritten image data, and output the rewritten image data.

[0264] In some embodiments, the controller 250 of the display device 200 described above can further perform the following steps:

[0265] The rewritten text content and the information representing the tone of the broadcast voice are input into a playing audio data processing model to obtain broadcast audio data corresponding to the rewritten image data; the tone information of the broadcast audio data is the same as the information representing the tone of the broadcast voice, and the text content corresponding to the broadcast audio data is the same as the rewritten text content; the playing audio data processing model is used to convert the input rewritten text content into corresponding initial broadcast audio data, perform adjustment processing on the initial broadcast audio data according to the information representing the tone of the broadcast voice, and output the broadcast audio data corresponding to the rewritten image data; the information representing the theme and the information representing the style are input into a background audio data processing model to obtain background audio data corresponding to the rewritten image data; the background audio label of the background audio data is a background audio label corresponding to the information representing the theme and the information representing the style; the background audio data processing model is used to perform background audio data generation processing according to the background audio label corresponding to the input information representing the theme and the information representing the style, and output the background audio data corresponding to the rewritten image data.

[0266] In some embodiments, as shown in Figure 1 and 2 The present application also provides a display device 200, which can include a display 260 and a controller 250; wherein the controller 250 is coupled with the display 260. As shown in Figure 9 The controller 250 is configured to:

[0267] Step S901, in the case that a picture book is displayed on a user interface, receiving a continuation demand information for the picture book, and taking an identification of a current picture book as a picture book identification corresponding to the continuation demand information.

[0268] Step S902, identifying text information corresponding to the continuation demand information, and acquiring picture book information corresponding to the picture book identification from a database according to the picture book identification corresponding to the continuation demand information; the picture book information represents multi-dimensional information of a picture book corresponding to the picture book identification.

[0269] Step S903, inputting the text information and information representing a role in the picture book information into a text processing model to obtain a continuation text content; the continuation text content represents a text content composed of the text information and the information representing the role; the text processing model is used to perform a picture book text content generation processing according to the input text information and the information representing the role, and output the continuation text content.

[0270] In step S904, the continuation text content, the information representing the style in the picture book information, and the information representing the character features in the picture book information are input into an image processing model to obtain continuation image data; the continuation image data represents image data matched with the continuation text content, the information representing the style, and the information representing the character features; the image processing model is used to perform picture book image data generation processing according to the input continuation text content, the information representing the style, and the information representing the character features, and output the continuation image data.

[0271] In step S905, based on the continuation text content and the information representing the voice in the picture book information, a play audio data synthesis processing is performed to obtain play audio data corresponding to the continuation image data, and based on the information representing the theme and the information representing the style in the picture book information, a background audio data generation processing is performed to obtain background audio data corresponding to the continuation image data.

[0272] In step S906, semantic recognition processing is performed on the text information corresponding to the continuation demand information to obtain a continuation position of the continuation image data in a picture book corresponding to a picture book identifier; the continuation image data is added in the continuation position in the picture book corresponding to the picture book identifier, and the play audio data and the background audio data corresponding to the continuation image data are configured to obtain a continuation picture book; the number of pages of the continuation picture book is greater than the number of pages of the picture book corresponding to the picture book identifier.

[0273] It should be noted that the specific implementation process of the display device 200 described above can refer to the related embodiments of the picture book continuation method shown in Figure 9 The specific implementation process of the display device 200 described above can refer to the related embodiments of the picture book continuation method shown in

[0274] The technical scheme provided in this embodiment is advantageous in that the multi-dimensional information of the original picture book is acquired and used as a constraint condition for continuation, which helps to ensure that the continuation content is consistent with the original picture book in terms of character image, artistic style, and story theme; the text content, image data, play audio, and background audio are collaboratively generated, and the appropriate continuation position is determined according to semantic recognition, which helps to realize natural connection of multi-modal content, expand the content of the picture book while maintaining the coherence and integrity of the story; thereby, the automatic continuation of the picture book is realized, the manual operation steps are reduced, and the continuation efficiency of the picture book is improved.

[0275] In some embodiments, the controller 250 of the display device 200 described above can further perform the following steps:

[0276] In the case of displaying a picture book on the user interface, a received input picture book selection operation is received; in response to the picture book selection operation, a picture book corresponding to a current focus is identified; the picture book corresponding to the current focus is identified as the selected picture book; in the case of identifying the selected picture book according to the picture book selection operation, continuation demand information for the picture book is received.

[0277] In some embodiments, the controller 250 of the display device 200 described above can also perform the following steps:

[0278] In the case where the continuation position is the end of the picture book corresponding to the picture book identifier, the continuation image data is added after the last page in the picture book corresponding to the picture book identifier, and the corresponding narration audio data and background audio data of the continuation image data are configured to obtain a picture book after continuation. In the case where the continuation position is the beginning of the picture book corresponding to the picture book identifier, the continuation image data is added before the first page in the picture book corresponding to the picture book identifier, and the corresponding narration audio data and background audio data of the continuation image data are configured to obtain a picture book after continuation.

[0279] In some embodiments, as shown in Figure 11 In order to more clearly describe the signaling interaction process between the modules of the display device 200, the present application further provides another picture book rewriting method, which can include the following steps:

[0280] Step 1: The picture book control in the control center receives user input content.

[0281] Step 2: The picture book control initiates a data pull request (including a character information pull request, a story theme pull request, a style information pull request, and a narration tone id pull request) to the database.

[0282] Step 3: The database returns character information, story theme, style information, character reference features, and narration tone id (identifier) to the picture book control.

[0283] Step 4: The picture book control sends the user input content and the character information obtained from the database to the story rewriting / continuation module in the multi-modal generation large model.

[0284] Step 5: The story rewriting / continuation module processes the received information to generate corresponding story content.

[0285] Step 6: The story rewriting / continuation module sends the generated story content to the text-to-image rewriting module, and can also send the story content to the picture book control.

[0286] Step 7: The text-to-image rewriting module converts the received story content into suitable prompt words for image generation.

[0287] Step 8: The text-to-image rewriting module sends the converted prompt words to the text-to-image module.

[0288] Step 9: The text-to-image module generates a composite image according to the prompt words (where style model switching can be performed), and returns the composite image to the picture book control.

[0289] Step 10, the voice synthesis module receives the broadcast voice id sent by the picture book control center, generates a voice broadcast and returns to the picture book control center.

[0290] Step 11, the music generation module receives the story theme information and style information sent by the picture book control center, generates background music based on the story theme information and style information, and returns to the picture book control center.

[0291] Step 12, the picture book control center stores the key information (database updates story content), and returns the image / broadcast audio / background audio to the user.

[0292] The technical scheme provided by the embodiment stores the picture book identifier and the multi-dimensional information in the database in association, and uniformly obtains the role, style, theme and other information based on the picture book identifier when rewriting, which is beneficial to avoid redundant operations such as repeated acquisition and manual input of information; the obtained information is respectively input into the text processing model, the image processing model and the audio processing model for automatic processing, and accurate replacement is performed while keeping the number of pages unchanged, which is beneficial to keep the consistency of the rewritten content and the original picture book in terms of role, style, theme and the like; thereby, it is beneficial to realize automatic rewriting of the picture book, reduce manual operation steps, and thereby improve the rewriting efficiency of the picture book.

[0293] In some embodiments, in order to more clearly illustrate the picture book rewriting method provided by the embodiments of the present application, the picture book rewriting method is specifically described below with a specific embodiment. The present application also provides another picture book rewriting method.

[0294] To meet the individual needs of users, the user needs to be given the right to edit and rewrite the picture book. For example, when the user feels that the generation of the 9th sentence of the picture book content does not meet his expectations, he can rewrite the story according to his needs, and in addition, the user is allowed to continue the story according to the plot.

[0295] From the perspective of the interaction link, in the TV scenario, this function needs voice recognition and large language model to understand the user's intention, and then the rewriting or continuation request is sent to the generation service.

[0296] For the generation service, it is necessary to ensure that the current rewritten or continued picture keeps the same role and style as the original story. The algorithm can use conditional control scheme or Train Free (free training) inter-frame Cross Attention (cross attention) scheme. This task requires the requester to provide not only the current continuation or rewriting content, but also the story context information or the condition information for controlling the generation. One way is to save the story generation dependency to the database system and assign a unique identifier to the story, which allows the user to edit and rewrite the story after a period of time.

[0297] The difficulty for the text-to-image model in the continuation and rewriting tasks is to maintain the consistency of the generated character attributes. This part involves Train Free or conditional control technology:

[0298] (1) Train Free, such as selecting similar story diffusion links, requires saving the feature input of different characters in different steps through the Attention layer in the U-shaped network during the generation of the previous story ij ∈F, where i traverses 1-N and j traverses 1-M, N is the number of generation steps, and M is the number of model layers. The output of the current step and the current layer will be referenced to f ij Cross Attention is done, which ensures the overall consistency of the output image with the story.

[0299] (2) Conditional control technology, such as Instance ID, randomly selects a picture containing the current rewriting character image in the story as a condition to control the rewriting of the picture book generation.

[0300] The user-operable position of the picture book rewriting can be referred to Figure 7 , the user clicks the remote control and focuses on a certain picture, at which point the voice input such as "change the story content to 'the little friend cries loudly and knocks the medicine bottle to the ground'".

[0301] The picture book application (APK, Android Package) sends the unique identifier of the story and the corresponding modified content to the picture book control center.

[0302] The picture book continuation trigger position can be referred to Figure 10 , the user controls the focus on a certain story through the remote control, and then inputs the voice such as "the story ends, the medicine pill invents a magical medicine, which makes the little friend never get sick".

[0303] The picture book application sends the unique identifier of the story and the user's request to the picture book control center.

[0304] The following is a module component in the timing diagram without continuation and rewriting:

[0305] 1. APP (application) control module: user interaction interface, including perception (sound button, text input window) and expression components (audio and video alignment, display and broadcast).

[0306] 2. Speech reception: access the sound microphone to receive user audio.

[0307] 3. Speech recognition: speech recognition module, converts speech to text.

[0308] 4. Compliance detection: Large language model, identify if user input contains sensitive words, if yes, refuse to generate and give user prompt.

[0309] 5. Role extraction: Large language model, according to user input, give appropriate role and role attributes, if user does not provide role, generate corresponding role and attributes according to story theme.

[0310] 6. Theme extraction: Large language model, according to user request, give appropriate story theme, overall theme style is positive and upward.

[0311] 7. Style extraction: Large language model, according to user input and story theme, specify corresponding picture book style, such as "cartoon", "realistic", "hand-drawn", "watercolor", etc., to make it easier to perform the story in the picture generation process.

[0312] 8. Voice tone recommendation: Large language model, default is "little girl", if the story theme is more brave and upward, "little boy" can be recommended, and other voice tone recommendations can also be given according to the story theme (system integrated voice tone), such as "mature male" and "mature female".

[0313] 9. Story generation: Large language model, according to user input, generate story content and story chapters (usually divided into 10-15 pages, each page of story control in 20 words).

[0314] 10. Text-to-image rewriting: Large language model, rewrite each page of story generated by story generation into prompt words suitable for image generation, such as "children are very happy" rewritten as "children hold hands up, face full of smile".

[0315] 11. Text-to-image: Text-to-image large model, according to text-to-image rewriting prompt words and recommended style, generate pictures.

[0316] 12. Speech synthesis: Multi-modal generation large model, convert story generation content into speech broadcast.

[0317] 13. Music generation: Multi-modal generation large model, according to recommended style and story theme, generate appropriate background music.

[0318] When including continuation and rewriting, the following modules need to be added:

[0319] 1. Picture book control: receive user demand from application program, issue to domain intent recognition module, according to intent classification result, decide whether to call database data, collect each sub-task return result, issue role information, story theme, story content, etc. to application program, store information such as role information, story theme, story content in database system.

[0320] 2. Storybook rewriting and sequel writing: Due to the general casualness of user input, key information needs to be extracted. This module is used for user key information extraction, and then the extracted information is integrated with the pulled database information to output the overall content of the story.

[0321] 3. Database: used to store role information, style information, story theme, story content, broadcast tone identification, and role reference characteristics (used to maintain the consistency of generated roles, which are role attribute-related characteristics extracted from the text-to-image model level).

[0322] The specific process can be referred to Figure 11 , Figure 12 and Figure 13 .

[0323] Figure 11 The information flow of the rewriting and sequel writing branch graph is shown, which can be referred to the description of Figure 11 above, and will not be repeated here.

[0324] Figure 12 The information flow of the overall architecture diagram is shown, and the information flow is as follows: the user opens the APP, the APP control module calls the voice listening, the user sends the far / near field voice to the voice receiving module in the information perception module, the voice receiving module converts the voice into text and sends it to the voice recognition module, the voice recognition module updates the text input area of the APP control module, the APP control module sends the user's input content to the picture book control in the control center, if the picture book control detects irregular content, it will feedback the irregular state to the APP control module, if the picture book control detects compliance and judges to be normal generation, it will be processed by the main branch of the story generation branch, if it is judged to be rewriting / sequel writing, it will be processed by the rewriting and sequel writing branch, the subsequent main branch or rewriting and sequel writing branch will return the image / broadcast audio / background audio to the picture book control, the picture book control will return the image / broadcast audio / background audio to the APP control module, and the APP control module will perform audio and video alignment and picture book broadcast synthesis processing.

[0325] Figure 13The information flow of the main branch diagram is shown: the content input by the user is first sent to the picture book control center in the control center, and the key information is stored in the database. The language model key information extraction / filling module includes role extraction, theme extraction, style extraction, and broadcast tone recommendation modules. The role extraction module generates role name and role attribute information; the theme extraction module generates story theme information; the style extraction module generates recommended style; the broadcast tone recommendation module generates recommended tone according to the content input by the user, and returns the broadcast tone information to the picture book control center. These information is sent to the multi-modal generation large model. The story generation module sends the story content to the text-to-image rewriting module, which converts the story content into suitable image generation prompts and sends them to the text-to-image module. The text-to-image module (which can switch between style models) generates a composite image according to the prompts and recommended style and returns it to the picture book control center. It can also return the role reference features to the picture book control center. The speech synthesis module receives the recommended tone and generates a synthesized speech broadcast that is returned to the picture book control center. The music generation module generates background music according to the recommended style and story theme information and returns it to the picture book control center. Among them, the database is also responsible for generating unique identifiers, storing role information, storing theme information, storing style information, storing tone information, storing story content, and storing role reference features. Finally, the picture book control center returns the image, broadcast audio, and background audio to the user.

[0326] Among them, the model level: the main scheme to maintain the consistency of the role is that for a certain role, a frame is randomly selected as a reference frame during the text-to-image process, and all other generated frames of the same role refer to the generation process characteristics of this frame. The reference method is cross attention mechanism.

[0327] The following is a systematic introduction to the model cross attention mechanism:

[0328] The consistency attention mechanism is used to solve the problem of maintaining the consistency of the role within the image batch. This attention mechanism does not need to be trained and can be plugged into the U-Net diffusion model. When used, the consistency attention is directly inserted into the original attention position of the U-Net architecture in the diffusion model, and the original self-attention weight is reused, so there is no need for training.

[0329] Given a batch of image features I∈R B×N×C , where B, N and C are the batch size, the number of tokens in each image, and the number of channels, respectively. Define a function Attention(X q ,X k ,X v ) to calculate self-attention. Xq, X k , X v represent the query Q, key K and value V used in attention calculation. The original self-attention is calculated as follows: iThe features I i are projected to Q i , K i , V i and fed into the attention function to get:

[0330]

[0331] To establish interaction between images within a batch to maintain consistency of the theme, consistency attention samples some reference frame features tokens from other images in the batch (save all generated image features before).

[0332] S i = RandSample(I1, I2, … In) i-1 , I i+1 , I B-1 , I B )

[0333] Where RandSample denotes the random sampling function. After sampling, the sampled tokens (reference frame features ref) are paired with the image features I i to form a new tokens set P i . Then P i is linearly projected to generate new keys and consistency attention values Here, the original Q i query does not change. Finally, the self-attention is calculated as follows:

[0334]

[0335]

[0336]

[0337] Where, is the weight of the reference image features and image features, K ref , V ref are the keys and values of the reference features.

[0338] Given the paired tokens, self-attention is performed in the image batch, thereby facilitating interaction between different image features. This type of interaction promotes convergence of the model on characters, faces, and clothing during the generation process. Although a simple and training-free approach is adopted, the consistency self-attention can effectively maintain the consistency of the main character.

[0339] For example, in terms of technical implementation, the scheme adopts a text-to-image diffusion model architecture with a consistency self-attention mechanism. A consistency attention module is added in the down-sampling and up-sampling layers of the diffusion U-shaped network. The structure of the consistency self-attention module includes a linear layer and corresponding dimension information. The entire processing flow includes a connection module, a label sampling module, a multiplication operation module, a soft-max processing module, a cross-attention module, a normalization processing module, a consistency self-attention module, a residual network module, and a consistency layer attention module. These modules work together to ensure consistency in the generated content throughout the processing from input to output. In specific implementation, the consistency attention mechanism is used to solve the problem of maintaining role consistency within an image batch. This mechanism does not require training and can be plugged into the diffusion model. It is achieved by inserting a consistency attention in the original attention position and reusing the original self-attention weight. The system maintains theme consistency by establishing interaction between images in the batch, randomly sampling features from other images in the batch, and pairing these features with the features of the current image to form a new feature set. This approach promotes interaction between different image features, effectively maintaining consistency in characters, faces, and clothing during the generation process.

[0340] The above embodiments are advantageous in realizing automatic updating of picture book stories, reducing manual operation steps, and thus improving the updating efficiency of picture book stories.

[0341] It should be understood that, although each step in the flowchart involved in each embodiment as described above is displayed in sequence according to the direction of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise specified herein, there is no strict order limitation for the execution of these steps, and these steps can be executed in other orders. Moreover, at least part of the steps in the flowchart involved in each embodiment as described above can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or steps or stages in other steps.

[0342] Based on the same inventive concept, the embodiments of the present application also provide a picture book rewriting device for implementing the picture book rewriting method described above and a picture book continuation device for implementing the picture book continuation method. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, and therefore the specific limitations in one or more picture book rewriting device embodiments and picture book continuation device embodiments provided below can refer to the limitations of the picture book rewriting method and the picture book continuation method described above, which will not be repeated here.

[0343] In one exemplary embodiment, asFigure 14 As shown in FIG. 1, a picture book rewriting device 1400 is provided, which can be applied to a picture book rewriting system 100 as shown in FIG. 1, and the picture book rewriting device 1400 can be applied to a picture book rewriting system 100 as shown in FIG. 1 Figure 1 As shown in FIG. 2, a display device 200 is provided, which can be applied to a picture book rewriting system 100 as shown in FIG. 1, and the display device 200 can be applied to a picture book rewriting system 100 as shown in FIG. 1 Figure 2 As shown in FIG. 2, the display device 200 can include a display 260 and a controller 250; the controller 250 is coupled with the display 260; the device can include:

[0344] A first receiving module 1401 is configured to receive rewriting requirement information for picture book content in a case where picture book content is displayed on a user interface, and take an identifier of a current picture book as a picture book identifier corresponding to the rewriting requirement information.

[0345] A first obtaining module 1402 is configured to identify text information corresponding to the rewriting requirement information, and obtain picture book information corresponding to the picture book identifier from a database according to the picture book identifier corresponding to the rewriting requirement information; the picture book information represents multi-dimensional information of a picture book corresponding to the picture book identifier.

[0346] A first input module 1403 is configured to input information representing a role in the text information and the picture book information into a text processing model to obtain rewritten text content; the rewritten text content represents text content composed of the text information and the information representing the role; the text processing model is used for performing picture book text content generation processing according to the input text information and the information representing the role, and outputting the rewritten text content.

[0347] A second input module 1404 is configured to input the rewritten text content, information representing a style in the picture book information, and information representing a role feature in the picture book information into an image processing model to obtain rewritten image data; the rewritten image data represents image data matched with the rewritten text content, the information representing the style, and the information representing the role feature; the image processing model is used for performing picture book image data generation processing according to the input rewritten text content, the information representing the style, and the information representing the role feature, and outputting the rewritten image data.

[0348] A first synthesizing module 1405 is configured to perform playing audio data synthesis processing based on the rewritten text content and information representing a playing voice in the picture book information to obtain playing audio data corresponding to the rewritten image data, and perform background audio data generation processing based on information representing a theme and information representing a style in the picture book information to obtain background audio data corresponding to the rewritten image data.

[0349] A data replacing module 1406 is configured to replace picture book image data in picture book content in a picture book corresponding to the picture book identifier with the rewritten image data, and replace playing audio data and background audio data corresponding to the picture book image data with playing audio data and background audio data corresponding to the rewritten image data to obtain a rewritten picture book; the number of pages of the rewritten picture book and the number of pages of the picture book corresponding to the picture book identifier remain unchanged.

[0350] In an example embodiment, as shown in Figure 15 a picture book continuation writing device 1500 is provided, which can be applied to a display device 200 as shown in Figure 1 The display device 200 can include a display 260 and a controller 250, the controller 250 being coupled to the display 260; the device can include: Figure 2

[0351] The second receiving module 1501 is configured to receive continuation writing demand information for a picture book in a case where a user interface displays a picture book, and take an identifier of a current picture book as a picture book identifier corresponding to the continuation writing demand information.

[0352] The second obtaining module 1502 is configured to identify text information corresponding to the continuation writing demand information, and obtain picture book information corresponding to the picture book identifier from a database according to the picture book identifier corresponding to the continuation writing demand information; the picture book information represents multi-dimensional information of a picture book corresponding to the picture book identifier.

[0353] The third input module 1503 is configured to input the text information and information representing a character in the picture book information into a text processing model to obtain continuation text content; the continuation text content represents text content composed of the text information and the information representing the character; the text processing model is configured to perform picture book text content generation processing according to the input text information and the information representing the character, and output the continuation text content.

[0354] The fourth input module 1504 is configured to input the continuation text content, information representing a style in the picture book information, and information representing a character feature in the picture book information into an image processing model to obtain continuation image data; the continuation image data represents image data matched with the continuation text content, the information representing the style, and the information representing the character feature; the image processing model is configured to perform picture book image data generation processing according to the input continuation text content, the information representing the style, and the information representing the character feature, and output the continuation image data.

[0355] The second synthesis module 1505 is configured to perform play audio data synthesis processing based on the continuation text content and information representing a play voice in the picture book information to obtain play audio data corresponding to the continuation image data, and perform background audio data generation processing based on information representing a theme and information representing a style in the picture book information to obtain background audio data corresponding to the continuation image data.

[0356] ​The data adding module 1506 is configured to perform semantic recognition processing on the text information corresponding to the continuation demand information, to obtain a continuation position of the continuation image data in the picture book corresponding to the picture book identifier; add the continuation image data in the continuation position in the picture book corresponding to the picture book identifier, and configure the broadcast audio data and the background audio data corresponding to the continuation image data, to obtain a continuation picture book; and the number of pages of the continuation picture book is greater than the number of pages of the picture book corresponding to the picture book identifier.

[0357] The modules in the picture book rewriting device and the picture book continuation device can be implemented by software, hardware, or a combination thereof. The modules can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in a computer device in software form, so as to be called and executed by a processor to perform operations corresponding to the modules.

[0358] In some embodiments, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program. The computer program is executed by a processor to implement the steps in the above method embodiments.

[0359] In some embodiments, a computer program product is provided, and the computer program product includes a computer program. The computer program is executed by a processor to implement the steps in the above method embodiments.

[0360] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use, and processing of the related data need to comply with relevant regulations.

[0361] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. The non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. The volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., without being limited thereto.

[0362] The technical features of the above embodiments can be combined in any manner. To make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combinations of the technical features do not exist contradictions, they should be considered as the scope of the present application.

[0363] The above-described embodiments are merely illustrative of several embodiments of the present application, which are described in more detail and in a specific manner, but should not be construed as limiting the scope of the patent of the present application. It should be noted that, for those of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A method for rewriting picture books, characterized in that, Applied to display devices, including: When the picture book content is displayed on the user interface, the system receives rewriting request information for the picture book content and uses the current identifier of the picture book as the picture book identifier corresponding to the rewriting request information. The text information corresponding to the rewriting requirement information is identified, and the picture book information corresponding to the picture book identifier is retrieved from the database according to the picture book identifier corresponding to the rewriting requirement information; the picture book information represents the multi-dimensional information of the picture book corresponding to the picture book identifier. The text information and the information representing the characters in the picture book information are input into the text processing model to obtain rewritten text content; the rewritten text content represents the text content composed of the text information and the information representing the characters; the text processing model is used to perform picture book text content generation processing based on the input text information and the information representing the characters, and output the rewritten text content. The rewritten text content, the style information in the picture book information, and the character feature information in the picture book information are input into an image processing model to obtain rewritten image data; the rewritten image data represents image data that matches the rewritten text content, the style information, and the character feature information; the image processing model is used to perform picture book image data generation processing based on the input rewritten text content, the style information, and the character feature information, and output the rewritten image data; Based on the rewritten text content and the information representing the broadcast timbre in the picture book information, audio data synthesis processing is performed to obtain broadcast audio data corresponding to the rewritten image data. Based on the information representing the theme and the information representing the style in the picture book information, background audio data generation processing is performed to obtain background audio data corresponding to the rewritten image data. The picture book image data in the picture book content corresponding to the picture book identifier is replaced with the rewritten image data, and the broadcast audio data and background audio data corresponding to the picture book image data are replaced with the broadcast audio data and background audio data corresponding to the rewritten image data to obtain the rewritten picture book; the number of pages of the rewritten picture book and the number of pages of the picture book corresponding to the picture book identifier remain unchanged.

2. The method according to claim 1, characterized in that, When the picture book content is displayed on the user interface, receiving rewriting request information for the picture book content includes: When the picture book content is displayed on the user interface, the user receives input to select the picture book content. If the selected picture book content is identified based on the picture book content selection operation, rewriting request information for the picture book content is received.

3. The method according to claim 2, characterized in that, After receiving the input of the picture book content selection, the process also includes: In response to the picture book content selection operation, identify the picture book content corresponding to the current focus; The picture book content corresponding to the current focus is identified as the selected picture book content.

4. The method according to claim 1, characterized in that, Before retrieving the picture book information corresponding to the picture book identifier from the database based on the picture book identifier corresponding to the rewritten requirement information, the process also includes: The text information corresponding to the rewrite requirement information is subjected to intent recognition processing to obtain the intent recognition result corresponding to the rewrite requirement information; the intent recognition result is used to indicate whether data is retrieved from the database. Based on the intent recognition result, determine whether to retrieve data from the database; The step of retrieving picture book information corresponding to the picture book identifier from the database based on the picture book identifier corresponding to the rewrite requirement information includes: If it is determined that data is to be obtained from the database, the picture book information corresponding to the picture book identifier is obtained from the database according to the picture book identifier corresponding to the rewrite requirement information.

5. The method according to claim 1, characterized in that, The information representing the character traits is also obtained through the following methods: From the picture books corresponding to the picture book identifier, select picture book image data containing the information representing the character as reference picture book image data; The information representing the character's features is extracted from the reference picture book image data.

6. The method according to claim 1, characterized in that, The step of inputting the text information and the information representing the characters in the picture book information into the text processing model to obtain the rewritten text content includes: The text information is input into a text processing model to obtain key text information; the key text information represents the key information in the text information; the text processing model is used to output the key text information in the text information based on the input text information. The key text information and the information representing the character are input into the text processing model to obtain the rewritten text content; the rewritten text content represents text content composed of the key text information and the information representing the character; the text processing model is used to perform picture book text content generation processing based on the input key text information and the information representing the character, and output the rewritten text content.

7. The method according to claim 1, characterized in that, The process of inputting the rewritten text content, the style information from the picture book information, and the character feature information from the picture book information into the image processing model to obtain rewritten image data includes: The rewritten text content is input into the text processing model to obtain image generation prompt information corresponding to the rewritten text content; the image generation prompt information is used to characterize text content adapted to the picture book image data generation and processing; the text processing model is used to adjust the input rewritten text content and output the adjusted text content; the adjusted text content is the image generation prompt information. The image generation prompt information, the representation style information, and the representation character feature information are input into the image processing model to obtain rewritten image data; the rewritten image data represents image data that matches the image generation prompt information, the representation style information, and the representation character feature information; the image processing model is used to perform picture book image data generation processing based on the input image generation prompt information, the representation style information, and the representation character feature information, and output the rewritten image data.

8. The method according to claim 7, characterized in that, The step of inputting the image generation prompt information, the representation style information, and the representation role feature information into the image processing model to obtain rewritten image data includes: The image generation prompt information and the information representing the character features are input into the image processing model to obtain initial rewritten image data; the initial rewritten image data represents image data that matches the image generation prompt information and the information representing the character features; the image processing model is used to perform picture book image data generation processing based on the input image generation prompt information and the information representing the character features, and output the initial rewritten image data; The initial rewritten image data and the representation style information are input into the image processing model to obtain rewritten image data; the style information of the rewritten image data is the same as the representation style information, and the image content of the rewritten image data is the same as the image content of the initial rewritten image data; the image processing model is used to render the representation style information of the input initial rewritten image data and output the rewritten image data.

9. The method according to claim 1, characterized in that, The process involves synthesizing playback audio data based on the rewritten text content and the information representing the broadcast timbre in the picture book information to obtain broadcast audio data corresponding to the rewritten image data. It also involves generating background audio data based on the information representing the theme and the style in the picture book information to obtain background audio data corresponding to the rewritten image data. This includes: The rewritten text content and the information representing the broadcast timbre are input into the playback audio data processing model to obtain broadcast audio data corresponding to the rewritten image data; the timbre information of the broadcast audio data is the same as the information representing the broadcast timbre, and the text content corresponding to the broadcast audio data is the same as the rewritten text content; the playback audio data processing model is used to convert the input rewritten text content into corresponding initial broadcast audio data, adjust the initial broadcast audio data according to the information representing the broadcast timbre, and output broadcast audio data corresponding to the rewritten image data; The information of the representation theme and the information of the representation style are input into the background audio data processing model to obtain background audio data corresponding to the rewritten image data; the background audio tags of the background audio data are background audio tags corresponding to the information of the representation theme and the information of the representation style; the background audio data processing model is used to perform background audio data generation processing based on the background audio tags corresponding to the input information of the representation theme and the information of the representation style, and output the background audio data corresponding to the rewritten image data.

10. A method for continuing a picture book, characterized in that, Applied to display devices, including: When a picture book is displayed on the user interface, the system receives continuation request information for the picture book and uses the current identifier of the picture book as the picture book identifier corresponding to the continuation request information. The text information corresponding to the continuation request information is identified, and the picture book information corresponding to the picture book identifier is retrieved from the database according to the picture book identifier corresponding to the continuation request information; the picture book information represents the multi-dimensional information of the picture book corresponding to the picture book identifier. The text information and the information representing the characters in the picture book information are input into the text processing model to obtain the continuation text content; the continuation text content represents the text content composed of the text information and the information representing the characters; the text processing model is used to perform picture book text content generation processing based on the input text information and the information representing the characters, and output the continuation text content. The continuation text content, the style information in the picture book information, and the character feature information in the picture book information are input into an image processing model to obtain continuation image data; the continuation image data represents image data that matches the continuation text content, the style information, and the character feature information; the image processing model is used to perform picture book image data generation processing based on the input continuation text content, the style information, and the character feature information, and output the continuation image data; Based on the continuation text content and the information representing the broadcast tone in the picture book information, playback audio data synthesis processing is performed to obtain broadcast audio data corresponding to the continuation image data; and based on the information representing the theme and the information representing the style in the picture book information, background audio data generation processing is performed to obtain background audio data corresponding to the continuation image data. Semantic recognition processing is performed on the text information corresponding to the continuation request information to obtain the continuation position of the continuation image data in the picture book corresponding to the picture book identifier; the continuation image data is added to the continuation position in the picture book corresponding to the picture book identifier, and the broadcast audio data and background audio data corresponding to the continuation image data are configured to obtain the continuation picture book; the number of pages of the continuation picture book is greater than the number of pages of the picture book corresponding to the picture book identifier.

11. The method according to claim 10, characterized in that, When the picture book is displayed on the user interface, receiving continuation request information for the picture book includes: When a picture book is displayed on the user interface, the user receives input to select a picture book. In response to the picture book selection operation, identify the picture book corresponding to the current focus; The picture book corresponding to the current focus is identified as the selected picture book; If the selected picture book is identified based on the picture book selection operation, the continuation request information for the picture book is received.

12. The method according to claim 10, characterized in that, The step of adding the continuation image data to the continuation position in the picture book corresponding to the picture book identifier, and configuring the broadcast audio data and background audio data corresponding to the continuation image data to obtain the continued picture book includes: When the continuation position is the end of the picture book corresponding to the picture book identifier, the continuation image data is added after the last page of the picture book corresponding to the picture book identifier, and the broadcast audio data and background audio data corresponding to the continuation image data are configured to obtain the continuation picture book; When the continuation position is at the beginning of the picture book corresponding to the picture book identifier, the continuation image data is added before the first page of the picture book corresponding to the picture book identifier, and the broadcast audio data and background audio data corresponding to the continuation image data are configured to obtain the continuation picture book.

13. A display device, characterized in that, include: monitor; The controller, coupled to the display, is configured to: When the picture book content is displayed on the user interface, the system receives rewriting request information for the picture book content and uses the current identifier of the picture book as the picture book identifier corresponding to the rewriting request information. The text information corresponding to the rewriting requirement information is identified, and the picture book information corresponding to the picture book identifier is retrieved from the database according to the picture book identifier corresponding to the rewriting requirement information; the picture book information represents the multi-dimensional information of the picture book corresponding to the picture book identifier. The text information and the information representing the characters in the picture book information are input into the text processing model to obtain rewritten text content; the rewritten text content represents the text content composed of the text information and the information representing the characters; the text processing model is used to perform picture book text content generation processing based on the input text information and the information representing the characters, and output the rewritten text content. The rewritten text content, the style information in the picture book information, and the character feature information in the picture book information are input into an image processing model to obtain rewritten image data; the rewritten image data represents image data that matches the rewritten text content, the style information, and the character feature information; the image processing model is used to perform picture book image data generation processing based on the input rewritten text content, the style information, and the character feature information, and output the rewritten image data; Based on the rewritten text content and the information representing the broadcast timbre in the picture book information, audio data synthesis processing is performed to obtain broadcast audio data corresponding to the rewritten image data. Based on the information representing the theme and the information representing the style in the picture book information, background audio data generation processing is performed to obtain background audio data corresponding to the rewritten image data. The picture book image data in the picture book content corresponding to the picture book identifier is replaced with the rewritten image data, and the broadcast audio data and background audio data corresponding to the picture book image data are replaced with the broadcast audio data and background audio data corresponding to the rewritten image data to obtain the rewritten picture book; the number of pages of the rewritten picture book and the number of pages of the picture book corresponding to the picture book identifier remain unchanged.

14. A display device, characterized in that, include: monitor; The controller, coupled to the display, is configured to: When a picture book is displayed on the user interface, the system receives continuation request information for the picture book and uses the current identifier of the picture book as the picture book identifier corresponding to the continuation request information. The text information corresponding to the continuation request information is identified, and the picture book information corresponding to the picture book identifier is retrieved from the database according to the picture book identifier corresponding to the continuation request information; the picture book information represents the multi-dimensional information of the picture book corresponding to the picture book identifier. The text information and the information representing the characters in the picture book information are input into the text processing model to obtain the continuation text content; the continuation text content represents the text content composed of the text information and the information representing the characters; the text processing model is used to perform picture book text content generation processing based on the input text information and the information representing the characters, and output the continuation text content. The continuation text content, the style information in the picture book information, and the character feature information in the picture book information are input into an image processing model to obtain continuation image data; the continuation image data represents image data that matches the continuation text content, the style information, and the character feature information; the image processing model is used to perform picture book image data generation processing based on the input continuation text content, the style information, and the character feature information, and output the continuation image data; Based on the continuation text content and the information representing the broadcast tone in the picture book information, playback audio data synthesis processing is performed to obtain broadcast audio data corresponding to the continuation image data; and based on the information representing the theme and the information representing the style in the picture book information, background audio data generation processing is performed to obtain background audio data corresponding to the continuation image data. Semantic recognition processing is performed on the text information corresponding to the continuation request information to obtain the continuation position of the continuation image data in the picture book corresponding to the picture book identifier; the continuation image data is added to the continuation position in the picture book corresponding to the picture book identifier, and the broadcast audio data and background audio data corresponding to the continuation image data are configured to obtain the continuation picture book; the number of pages of the continuation picture book is greater than the number of pages of the picture book corresponding to the picture book identifier.

Citation Information

Patent Citations

  • Electronic picture book generation method and device and electronic equipment

    CN114693844A

  • Digital picture book identification method and system, electronic equipment and storage medium

    CN118885969A