Electronic picture book generation method and device, storage medium and computer equipment

Through a flexible and efficient method of generating electronic picture book, users can customize story content and interactively generate target images, solving the problem of time-consuming and labor-intensive electronic picture book production in the prior art, and achieving high flexibility and diversified picture book generation.

CN120067381APending Publication Date: 2025-05-30BEIJING YIZHEN XUESI EDUCATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510202480.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The production process of existing electronic picture books relies on manual creation by professional designers, which is time-consuming and labor-intensive, and it is difficult to quickly respond to market changes and personalized needs.

Method used

It provides a flexible and efficient method for generating electronic picture books, allowing users to customize story content according to their own needs and preferences, and generate target images that meet expectations through user interaction, thereby quickly generating electronic picture books that meet users' preferences.

Benefits of technology

It greatly improves the flexibility of adjusting the content and style of picture books, can meet the diverse needs of users and enhance the user's creative experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067381A_ABST
    Figure CN120067381A_ABST
Patent Text Reader

Abstract

The invention discloses an electronic picture book generation method and device, a storage medium and computer equipment, and the method comprises the steps: obtaining first user input data, obtaining a target story text according to the first user input data, and carrying out the text segmentation of the target story text, and obtaining a plurality of sub-texts; and for each sub-text, performing text analysis on the sub-text to obtain a target keyword, generating a basic image corresponding to the sub-text according to the target keyword, and determining a target image corresponding to the sub-text based on the basic image and second user input data, the second user input data is used for indicating an image confirmation operation or an image modification operation; and according to each sub-text and the corresponding target image, generating a picture book page corresponding to the sub-text, and according to the picture book page corresponding to each sub-text, generating a target electronic picture book.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a method and device for generating electronic picture books, a storage medium, and a computer device. Background Art

[0002] In the digital age, as an emerging media form, electronic picture books are gradually replacing traditional paper picture books and becoming an important part of the field of children's education and entertainment. Due to their rich visual expressiveness, interactivity, and portability, electronic picture books are widely loved by the public. However, the current production process of electronic picture books mostly relies on the manual creation of professional designers, which is not only time-consuming and laborious but also difficult to quickly respond to market changes and personalized needs.

[0003] The production process of traditional electronic picture books usually includes multiple links such as story writing, illustration design, and page layout, and each link requires the in-depth participation of professionals. Especially in the illustration design link, designers need to manually draw or select appropriate picture materials according to the plot and text content of the story. This process not only tests the artistic creation ability of designers but is also limited by their personal style and understanding of the story. In addition, with the diversification of user needs, how to flexibly adjust the content and style of picture books according to the preferences and feedback of specific user groups has become an urgent problem to be solved in the production of electronic picture books. Summary of the Invention

[0004] In view of this, this application provides a method and device for generating electronic picture books, a storage medium, and a computer device, providing a flexible and efficient method for generating electronic picture books. This method allows users to customize story content according to their own needs and preferences, modify the basic images generated by the system to obtain target images that meet user expectations, so as to quickly generate electronic picture books that meet user preferences, which can greatly improve the flexibility of adjusting the content and style of picture books, is conducive to meeting the diverse needs of users, and enhances the user's creative experience.

[0005] According to one aspect of this application, a method for generating an electronic picture book is provided, including:

[0006] Obtain first user input data, and based on the first user input data, obtain a target story text, and perform text segmentation on the target story text to obtain multiple sub-texts;

[0007] For each sub-text, perform text parsing on the sub-text to obtain target keywords, and based on the target keywords, generate a basic image corresponding to the sub-text, and determine a target image corresponding to the sub-text based on the basic image and second user input data, where the second user input data is used to indicate an image confirmation operation or an image modification operation;

[0008] According to each sub - text and the corresponding target image, generate the picture book page corresponding to the sub - text, and generate the target electronic picture book according to the picture book pages corresponding to each sub - text.

[0009] According to another aspect of the present application, there is provided a device for generating an electronic picture book, including:

[0010] A text segmentation module, configured to obtain first user input data, obtain a target story text according to the first user input data, and perform text segmentation on the target story text to obtain a plurality of sub - texts;

[0011] An image generation module, configured to, for each sub - text, perform text parsing on the sub - text to obtain target keywords, generate a basic image corresponding to the sub - text according to the target keywords, and determine the target image corresponding to the sub - text based on the basic image and second user input data, where the second user input data is used to indicate an image confirmation operation or an image modification operation;

[0012] A picture book generation module, configured to generate the picture book page corresponding to the sub - text according to each sub - text and the corresponding target image, and generate the target electronic picture book according to the picture book pages corresponding to each sub - text.

[0013] According to yet another aspect of the present application, there is provided a storage medium, on which a computer program is stored, and when the program is executed by a processor, the above - mentioned method for generating an electronic picture book is implemented.

[0014] According to still another aspect of the present application, there is provided a computer device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, and when the processor executes the program, the above - mentioned method for generating an electronic picture book is implemented.

[0015] With the above technical solution, a method and apparatus for generating an electronic picture book, a storage medium, and a computer device provided by this application first obtain a target story text based on the first user input data. After determining the target story text, the target story text can be segmented into multiple sub-texts. For each sub-text obtained after segmentation, each sub-text can be parsed separately to extract target keywords. Subsequently, based on the extracted target keywords, a basic image corresponding to the target keywords can be generated through a large model. After generating the basic image, interaction with the user is allowed to continue. At this time, the user can choose to directly confirm the basic image or modify the basic image to obtain a target image. Finally, each sub-text and its corresponding target image can be combined together to generate a picture book page, and all the picture book pages can be combined in sequence to form a complete electronic picture book. The embodiments of this application provide a flexible and efficient method for generating an electronic picture book. This method allows users to customize the story content according to their own needs and preferences, modify the basic image generated by the system to obtain a target image that meets the user's expectations, thereby quickly generating an electronic picture book that meets the user's preferences, which can greatly improve the flexibility of adjusting the content and style of the picture book, is conducive to meeting the diverse needs of users, and enhances the user's creative experience.

[0016] The above description is only an overview of the technical solution of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of this application more obvious and understandable, the specific embodiments of this application are specifically given below. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of this application and constitute a part of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0018] Figure 1 A flowchart showing a method for generating an electronic picture book provided by an embodiment of this application is shown;

[0019] Figure 2 A schematic structural diagram of an apparatus for generating an electronic picture book provided by an embodiment of this application is shown;

[0020] Figure 3 A schematic diagram of the device structure of a computer device provided by an embodiment of this application is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The following will refer to the drawings and combine the embodiments to detail this application. It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other.

[0022] In this embodiment, a method for generating an electronic picture book is provided. As Figure 1 shown, the method includes:

[0023] Step 101, obtain first user input data, obtain a target story text according to the first user input data, and perform text segmentation on the target story text to obtain multiple sub-texts.

[0024] A method for generating an electronic picture book provided by an embodiment of this application can automatically generate an electronic picture book that meets user expectations according to user input. Specifically, when a user wants to generate an electronic picture book, the user can input data through an application interface, a Web interface, etc. Then, the data input by the user can be obtained, which is named first user input data here, and based on the first user input data, a target story text can be obtained. Among them, the first user input data can be text, voice, pictures, etc., specifically depending on the input method designed by the system, or can include multiple methods at the same time, and the user can choose their preferred input method. When the first user input data is text or voice, the text or voice can be several keywords, and these keywords can be the main content of the story that the user wants to emphasize; the text or voice can also correspond to a complete story. When the first user input data is keywords, a complete target story text can be generated according to the keywords through natural language processing (NLP, Natural Language Processing) technology; when the first user input data is a complete story, the story can be directly used as the target story text. In another embodiment, when it is detected that the first user input data is a complete story, the text size of the text corresponding to the story can be judged. If the text size is less than a first preset text threshold, it can be expanded through natural language processing technology, and the text size after expansion can be greater than a second preset text threshold, and the story text after expansion is the target story text. Among them, the first preset text threshold and the second preset text threshold can be the same, or the second preset text threshold is greater than the first preset text threshold, so that the target story text can be more substantial. It should be noted that when the first user input data is voice data, the voice data can be first converted into text, and then the target story text can be obtained according to the text obtained after conversion. In addition, when the first user input data is a picture, the first user input data can be input into a large model, and the large model can describe the picture, and the obtained description text can be used as the target story text.

[0025] After determining the target story text, the target story text can be segmented into multiple sub-texts, and each sub-text can be a sentence, a paragraph, etc. Subsequently, corresponding picture book pages can be generated for each sub-text respectively.

[0026] Step 102: For each sub - text, perform text parsing on the sub - text to obtain target keywords. According to the target keywords, generate a basic image corresponding to the sub - text, and based on the basic image and the second user input data, determine the target image corresponding to the sub - text, where the second user input data is used to indicate an image confirmation operation or an image modification operation.

[0027] In this embodiment, for each sub - text obtained after segmentation, each sub - text can be parsed separately to extract target keywords. The target keywords can include character keywords, action keywords, scene keywords, item keywords, etc. NLP techniques such as semantic analysis and sentiment recognition can be used in the parsing process, which is not limited here. Subsequently, based on the extracted target keywords, a basic image corresponding to the target keywords can be generated through a large - model. However, sometimes the basic image generated by the large - model may not fully meet the user's needs, so a function for modifying the basic image is provided. Furthermore, after generating the basic image, interaction with the user is allowed to continue. At this time, the user can choose to directly confirm the basic image or modify the basic image. Specifically, user input data can be obtained again, named the second user input data here, and the target image corresponding to each sub - text is determined according to the second user input data. For example, if the second user input data is the confirmation data of the basic image, then the basic image can be directly used as the target image at this time; if the second user input data is the modification data of the basic image, then the basic image can be modified based on the user's modification operation, and the target image is obtained after the modification is completed.

[0028] Step 103: Generate a picture book page corresponding to the sub - text according to each sub - text and the corresponding target image, and generate a target e - picture book according to the picture book pages corresponding to each sub - text.

[0029] In this embodiment, each sub - text and its corresponding target image can be combined together to generate a picture book page. Subsequently, all the picture book pages are combined in sequence to form a complete e - picture book.

[0030] By applying the technical solution of this embodiment, first, based on the first user input data, a target story text is obtained. After determining the target story text, the target story text can be segmented into multiple sub-texts. For each sub-text obtained after segmentation, each sub-text can be parsed separately to extract target keywords. Subsequently, based on the extracted target keywords, a basic image corresponding to the target keywords can be generated through a large model. After generating the basic image, interaction with the user is allowed to continue. At this time, the user can choose to directly confirm the basic image or modify the basic image to obtain a target image. Finally, each sub-text and its corresponding target image can be combined together to generate a picture book page, and all the picture book pages can be combined in sequence to form a complete electronic picture book. The embodiment of the present application provides a flexible and efficient method for generating an electronic picture book. This method allows users to customize the story content according to their own needs and preferences, modify the basic image generated by the system to obtain a target image that meets the user's expectations, thereby quickly generating an electronic picture book that meets the user's preferences, greatly improving the flexibility of adjusting the content and style of the picture book, facilitating meeting the diverse needs of users, and enhancing the user's creative experience.

[0031] In the embodiment of the present application, optionally, the "generating the basic image corresponding to the sub-text according to the target keyword" in step 102 includes: generating a text-to-image prompt text corresponding to the sub-text according to the target keyword and a preset expansion rule, and inputting the text-to-image prompt text into the large model to obtain the basic image corresponding to the sub-text.

[0032] In this embodiment, the basic image corresponding to each sub-text can be generated in the following manner. First, a text-to-image prompt text corresponding to the sub-text is generated according to the target keyword and a preset expansion rule. Here, the preset expansion rule is a set of criteria for enriching and refining the target keyword, aiming to generate a more descriptive and detailed text-to-image prompt text. The text-to-image prompt text can be detailed and specific enough for the large model to accurately understand and generate an image that meets the expectations. In addition, the preset expansion rule can also include a rule for expanding the text describing the style of the image scene. For example, if the target keyword includes a scene keyword: forest, dusk, sunny day; a character keyword: little girl; and an item keyword: basket, bread, milk. Then the text-to-image prompt text generated according to these target keywords and the preset expansion rule can be: "A young girl is walking on a forest path at the golden sunset in a bright red hooded cloak. She is carrying a wicker basket filled with bread and milk in her hand. The forest is lush and peaceful, and the warm evening sun shines through the leaves, forming long shadows. The light in the picture is soft, and the environmental details are rich, presenting the style of a storybook illustration."

[0033] Large models generally refer to deep learning models with powerful generation capabilities, such as generative adversarial networks (GANs), variational autoencoders (VAEs), or Transformer-based generative models. These models are trained with a large amount of data and can generate high-quality images according to the input text-to-image prompt text. Therefore, after obtaining the text-to-image prompt text, the generated text-to-image prompt text can be input into the large model. Based on the input text-to-image prompt text, the large model can generate a base image corresponding to the sub-text content. The base images corresponding to different sub-texts can be generated simultaneously or successively. By combining the target keywords, preset expansion rules, and the generation capabilities of the large model, the embodiments of the present application not only improve the generation accuracy and detail richness of the base image, but also enhance the visual effect and storytelling of the e-picture book.

[0034] In an embodiment of the present application, optionally, for each sub - text in non - the first sub - text, the text - to - image prompt text further includes a guiding label constructed based on the context information corresponding to the sub - text. The context information includes the text - to - image prompt text corresponding to the previous sub - text of the current sub - text, or the target image corresponding to the previous sub - text of the current sub - text. The step of "generating the text - to - image prompt text corresponding to the sub - text according to the target keyword and the preset expansion rule, and inputting the text - to - image prompt text into the large - model to obtain the base image corresponding to the sub - text" includes: when the context information includes the text - to - image prompt text corresponding to the previous sub - text of the current sub - text, for the first sub - text, generating the text - to - image prompt text corresponding to the first sub - text according to the target keyword and the preset expansion rule corresponding to the first sub - text, inputting the text - to - image prompt text corresponding to the first sub - text into the large - model to obtain the base image corresponding to the first sub - text; for each sub - text in non - the first sub - text, using the text - to - image prompt text corresponding to the previous sub - text of the current sub - text as a guiding label, generating the text - to - image prompt text corresponding to the current sub - text according to the target keyword, the preset expansion rule, and the guiding label, and inputting the text - to - image prompt text corresponding to the current sub - text into the large - model to obtain the base image corresponding to the current sub - text; when the context information includes the target image corresponding to the previous sub - text of the current sub - text, for the first sub - text, generating the text - to - image prompt text corresponding to the first sub - text according to the target keyword and the preset expansion rule corresponding to the first sub - text, inputting the text - to - image prompt text corresponding to the first sub - text into the large - model to obtain the base image corresponding to the first sub - text; for each sub - text in non - the first sub - text, using the target image corresponding to the previous sub - text of the current sub - text as a guiding label, generating the text - to - image prompt text corresponding to the current sub - text according to the target keyword, the preset expansion rule, and the guiding label, and inputting the text - to - image prompt text corresponding to the current sub - text into the large - model to obtain the base image corresponding to the current sub - text.

[0035] In this embodiment, for non - the first sub - text, the text - to - image prompt text may further include a guiding label, and the guiding label may be constructed according to the context information of the current sub - text. Here, the context information may include two cases. One is the text - to - image prompt text corresponding to the previous sub - text of the current sub - text, and the other is the target image corresponding to the previous sub - text of the current sub - text. Through the guiding label, a connection can be established between the base image of the current sub - text generated by the large - model and the previous sub - texts, avoiding the irrelevance between each base image.

[0036] Specifically, when the context information is the text-to-image prompt text corresponding to the previous sub-text of the current sub-text, for the first sub-text: According to the target keyword of the first sub-text, apply the preset expansion rule to generate the text-to-image prompt text of the first sub-text. Then, this text-to-image prompt text can be input into the large model to generate the basic image corresponding to the first sub-text. For non-first sub-texts (the second sub-text to the last sub-text): First, obtain the text-to-image prompt text of its previous sub-text as the guiding label. Then, combine the target keyword, the preset expansion rule, and the guiding label to generate the text-to-image prompt text of the current sub-text. Here, the guiding label can help the large model maintain the consistency of the basic image sequence in terms of theme and style. Subsequently, the generated text-to-image prompt text can be input into the large model to generate the basic image corresponding to the current sub-text. It should be noted that when generating the text-to-image prompt text of the current sub-text according to the target keyword, the preset expansion rule, and the guiding label, the target keyword corresponding to the current sub-text can be expanded into a preliminary text-to-image prompt text through the preset expansion rule, and then the guiding label and the preliminary text-to-image prompt text are packaged to generate the text-to-image prompt text corresponding to the current sub-text. When the large model analyzes the text-to-image prompt text of the current sub-text, it can identify which part is the guiding label and which part is the initial text-to-image prompt text to distinguish the main dependent part (the initial text-to-image prompt text) for generating the basic image and the part for fine-tuning the image (the guiding label).

[0037] When the context information is the target image corresponding to the previous sub-text of the current sub-text, for the first sub-text: the process is the same as above, generating the text-to-image prompt text for the first sub-text and generating the base image. For non-first sub-texts: first obtain the target image of its previous sub-text as a guiding label. Then, combine the target keyword, the preset expansion rule, and the guiding label (here in the form of an image) to generate the text-to-image prompt text for the current sub-text. The image guiding label here provides a more intuitive visual context, helping the model generate a new image that is visually coherent with the previous image. After that, input the generated text-to-image prompt text into the large model to generate the base image corresponding to the current sub-text. It should be noted that similarly, when combining the target keyword, the preset expansion rule, and the guiding label to generate the text-to-image prompt text for the current sub-text, the target keyword corresponding to the current sub-text can be expanded into a preliminary text-to-image prompt text through the preset expansion rule, and then the guiding label and the preliminary text-to-image prompt text are packaged to generate the text-to-image prompt text corresponding to the current sub-text. Among them, the large model at this time is a model that can update the target image based on the text-to-image prompt text of the current sub-text, and the updated image is the base image corresponding to the current sub-text. Here, the guiding label can provide a more intuitive visual context, helping the large model generate a new base image that is visually coherent with the target image corresponding to the previous sub-text. At the same time, since the target image is an image that meets the user's needs, generating the base image of the current sub-text based on this image can further improve the user satisfaction of the base image, avoid the user making the same modifications to each base image, and greatly improve the generation efficiency and user experience of the electronic picture book.

[0038] In an embodiment of the present application, optionally, when the second user input data is used to indicate an image modification operation, in step 102, "determining the target image corresponding to the sub-text based on the base image and the second user input data" includes: determining an image modification type based on the second user input data, where the image modification type is one of element addition, element deletion, and element modification, and the element includes at least one of a scene element, a character element, and an item element; when the image modification type is element deletion, based on the user's area selection operation, determining a target selection area from the base image, identifying the element to be deleted included in the target selection area, and performing a deletion operation on the element to be deleted to obtain the target image corresponding to the sub-text; when the image modification type is element modification, based on the user's area selection operation, determining a target selection area from the base image, identifying the element to be replaced included in the target selection area, performing a deletion operation on the element to be replaced to obtain an updated base image, and integrating the new element into the base image to obtain the target image corresponding to the sub-text; when the image modification type is element addition, integrating the new element into the base image to obtain the target image corresponding to the sub-text.

[0039] In this embodiment, the base image is composed of multiple elements. Here, the elements can be scene elements (such as backgrounds, environments), character elements (such as people, animals), and item elements (such as props, decorations), etc. After obtaining the base image, the user can modify the base image generated by the large model. The modification methods include three types, namely element addition, element deletion, and element modification. Specifically, the user can select a region each time and modify the elements included in that region.

[0040] The system can determine the image modification type through the second user input data input by the user. If it is determined that the image modification type is element deletion, then at this time, the user can be required to perform an area selection operation to specify the position of the element to be deleted from the base image. The user can determine the target selection area through a certain interface tool (such as a mouse, a touchpad, etc.), and the system can identify the element to be deleted included in the target selection area. Then, a deletion operation is performed on the identified element to be deleted, so as to obtain an updated image, and this updated image is the target image corresponding to the current sub-text.

[0041] If it is determined that the image modification type is element modification, the user can be required to perform a region selection operation at this time to specify the position where the element to be replaced is located. After the user determines the target selection region, the system can identify the element to be replaced in the target selection region, and first perform a deletion operation on the element to be replaced to obtain an updated base image. Then, the new element (i.e., the element that the user hopes to replace with) is integrated into the updated base image, so as to obtain the target image corresponding to the current sub-text.

[0042] If it is determined that the image modification type is element addition, there is no need for the user to perform a region selection operation at this time, because the new element usually does not depend on a specific position existing in the base image. The system can directly integrate the new element into the base image according to the user's operation. This integration process can involve adjusting attributes such as the size, position, and transparency of the element to ensure that the new element is coordinated with other parts of the base image. Finally, the target image corresponding to the current sub-text is obtained.

[0043] It should be noted that the user can also repeatedly modify the same base image multiple times through the above method. The embodiment of the present application provides a method for modifying a base image, which improves the flexibility of picture book image generation and makes the finally generated picture book image closer to the user's needs and expectations.

[0044] In the embodiment of the present application, optionally, the "integrating the new element into the base image to obtain the target image corresponding to the sub-text" includes: determining the placement position and scaling parameter corresponding to the new element, and integrating the new element into the base image according to the placement position and scaling parameter corresponding to the new element to obtain an initial image; performing layout analysis on the initial image to obtain a layout score, where the layout score includes at least one of the placement position score of the new element, the visual center of gravity score of the initial image, and the color score of the initial image; when there is a score in the layout score that is less than a preset score, outputting a layout adjustment interaction button, and after receiving a trigger instruction of the layout adjustment interaction button, performing layout adjustment on the new element in the initial image, and using the adjusted initial image as the target image; when all scores in the layout score are greater than or equal to the preset score, using the initial image as the target image.

[0045] In this embodiment, regardless of whether the image modification type is element modification or element addition, the newly added elements can ultimately be integrated into the base image in the following manner to generate the target image of the sub-text. Specifically, first, according to the user's operation instructions, the placement position and scaling parameters of the newly added elements (which can be scene elements, character elements, item elements, etc.) in the base image can be determined. Both the placement position and the scaling parameters are controlled by the user. Then, according to the determined placement position and scaling parameters, the newly added elements are integrated into the base image to form an initial image. Next, layout analysis is performed on the integrated initial image to evaluate its overall visual effect. The layout score can include multiple aspects, such as: Placement position score of the newly added elements: Evaluate whether the position of the newly added elements in the initial image is appropriate and whether design principles such as alignment, contrast, and repetition are followed. Visual center of gravity score: Analyze whether the visual center of gravity of the initial image is balanced, whether it attracts the audience's attention, and whether it highlights important information. Color score: Evaluate whether the color combination of the initial image is harmonious and whether it conforms to the design theme or emotional expression. Among them, each scoring item has a preset score for judging whether the current score meets the requirements.

[0046] If any item in the layout score is lower than the preset score, it indicates that some aspects of the initial image need to be improved. At this time, a layout adjustment interaction button can be output to prompt the user to perform layout adjustment. When the user clicks the layout adjustment interaction button, the system receives the trigger instruction and allows the user to adjust the layout of the newly added elements in the initial image. It should be noted that after the user triggers the layout adjustment interaction button, the user can manually adjust the newly added elements, or the system can automatically calculate the optimal placement position and optimal scaling parameters of the newly added elements based on the placement position of the newly added elements, the visual center of gravity of the initial image, and the color of the initial image. After the user triggers the layout adjustment interaction button, the optimal adjustment of the layout can be achieved with one key without the need for the user to manually adjust. After manual adjustment, layout analysis can be performed again to obtain the layout score... If it is an automatic adjustment, then after one-key adjustment, the adjusted initial image is the target image. Alternatively, the user can also choose not to trigger after the layout adjustment interaction button is output and default to using the initial image with a layout score less than the preset score as the target image.

[0047] If all items in the layout score are greater than or equal to the preset score, it indicates that the layout of the initial image is already good enough and no further adjustment is required. At this time, the initial image is directly used as the target image. The embodiment of the present application ensures that the finally generated target image reaches the best state in terms of visual effect and meets the design requirements and user expectations through layout analysis and user interaction.

[0048] In an embodiment of the present application, optionally, "performing layout analysis on the initial image to obtain a layout score" includes: based on the placement position of the new element in the initial image, determining a target analysis area corresponding to the placement position, and identifying the original elements included in the target analysis area, determining an element intersection between the new element and the original elements according to the placement position and the target position corresponding to the original elements, calculating the area ratio of the element intersection to the area of the new element, and calculating the placement position score corresponding to the new element based on the area ratio; and / or, based on preset composition principles and / or user image preferences, determining the ideal center of gravity corresponding to the initial image, and according to the position, element size, and element color corresponding to each element in the initial image, determining the element weight corresponding to each element, calculating the weighted position corresponding to the element according to the element center point position and element weight of each element, adding up the weighted positions corresponding to each element to obtain the visual center of gravity corresponding to the initial image, and calculating the visual center of gravity score corresponding to the initial image based on the ideal center of gravity and the visual center of gravity; and / or, inputting the initial image into a preset color evaluation model, respectively extracting the contrast feature, color harmony feature, and color distribution feature of the initial image through the preset color evaluation model, performing feature fusion on the contrast feature, color harmony feature, and color distribution feature to obtain a target fusion feature, and calculating the color score corresponding to the initial image according to the target fusion feature.

[0049] In this embodiment, the placement position score, the visual center of gravity score, and the color score can be calculated respectively.

[0050] First, the placement position score can be determined based on the following method.

[0051] First, according to the placement position of the new element in the initial image, a target analysis area is determined. The target analysis area can be an area extending a certain distance around the placement position of the new element in the initial image. Here, the extended distance can be determined according to experience. Then, within the target analysis area, all the original elements are identified. Here, the original elements refer to the elements that originally existed in the initial image. Then, according to the placement position and shape of the new element, and the target position and shape of the original elements, the overlapping part between the new element and the original elements, that is, the element intersection, is determined. Calculate the ratio of the area of the element intersection to the total area of the new element. This ratio reflects the conflict degree or overlapping degree between the new element and the original elements. Finally, based on the area ratio, the placement position score is calculated through a set scoring criterion or function. For example, if the area ratio is small, it means that the conflict between the new element and the original elements is small, and a higher score can be given; otherwise, a lower score is given.

[0052] Second, the visual center of gravity score can be determined based on the following methods.

[0053] When designing an image, if the center of gravity of the image can be well grasped, the visual attraction of the image can be greatly improved. First, an ideal center of gravity can be determined from the initial image according to the preset composition principles (such as the rule of thirds, etc.) and / or the user's image preferences. The rule of thirds is a simple and effective composition technique that divides the image into nine equal parts, and attracts the audience's attention by placing important image elements at the four intersection points (or called "points of interest") formed by the nine dividing lines. Therefore, after determining the intersection points, it can be analyzed which intersection point the character elements in the initial image are near, and the intersection point closest to the character elements is used as a candidate for the ideal center of gravity of the initial image. If there are multiple character elements in the initial image at the same time, the intersection point surrounded by the most character elements can be used as the ideal center of gravity of the initial image. In addition, picture book images generated within a preset time period before the current time of the user can be obtained, the sizes of these picture book images can be unified (such as scaled to the same size), the visual center of gravity corresponding to these unified picture book images can be analyzed (which can be determined by existing methods), and the average value of the visual center of gravity corresponding to these picture book images is taken to obtain the ideal center of gravity corresponding to the user's image preferences. If both the preset composition principle and the user's image preferences are considered when determining the ideal center of gravity, different weights can be assigned to the ideal centers of gravity determined by the preset composition principle and the user's image preferences, and then the final ideal center of gravity is obtained by weighting. Here, the weight can be determined according to actual needs.

[0054] Second, according to the position, element size, and element color of each element in the initial image, an element weight is assigned to each element. The size of the element weight reflects the importance of the element in the initial image. Among them, the position of the element in the initial image determines its importance or salience. For example, an element located at the center of the initial image is considered more important and can be assigned a higher weight, while an element at the edge has a lower weight. In addition, larger-sized elements often contain more information or are considered more salient, so they can be assigned a higher weight. On the contrary, smaller elements contain less information or lower salience, so the weight is lower. Color can reflect the salience or emotional value of the element. For example, bright colors (such as red, yellow) are more eye-catching than dull colors (such as gray, blue), so bright colors are assigned a higher weight. In one embodiment, a weight ratio can also be assigned to each attribute (position, size, color) according to actual needs. These weight ratios can be determined based on experience, statistics, or expert opinions. For each element, after determining the weight of each attribute of the element, the weight of each attribute is multiplied by its corresponding weight ratio, and then these products are added together to obtain the element weight of each element.

[0055] Then, calculate the element center point position of each element, and perform weighted processing according to its element weight to obtain the weighted position. Add up the weighted positions of all elements to obtain the visual center of gravity position of the initial image. Finally, calculate the visual center of gravity score based on the distance or difference degree between the ideal center of gravity and the visual center of gravity. Here, if the visual center of gravity is close to the ideal center of gravity, a higher score is given; otherwise, a lower score is given.

[0056] Third, the color score can be determined based on the following method.

[0057] First, input the initial image into a preset color evaluation model, which can extract the contrast feature, color harmony feature, and color distribution feature of the image. Here, the preset color evaluation model can be a pre-trained model for extracting and evaluating the contrast feature, color harmony feature, and color distribution feature of the image. This model can be constructed based on machine learning, deep learning, or other algorithms. (1) Contrast refers to the degree of difference between the bright and dark regions in the image. In a high-contrast image, the bright and dark regions are distinct, and details are clearly visible; while in a low-contrast image, it may appear blurred or monotonous. For the extraction of this feature, specifically, the preset color evaluation model can analyze the pixel values of the image, calculate the brightness difference, and thus extract the contrast feature. (2) Color harmony refers to the degree of harmony of the color combination in the image. A harmonious color combination can produce a visual sense of pleasure, while a disharmonious color combination may make people feel uncomfortable. For the extraction of this feature, specifically, the preset color evaluation model can analyze the color distribution and combination relationship in the image, calculate the similarity and contrast between colors, and thus extract the color harmony feature. Algorithms such as color space conversion and color clustering analysis can be used in the extraction process. (3) Color distribution refers to the proportion and distribution of different colors in the image. Different color distributions can produce different visual effects and emotional experiences. For the extraction of this feature, specifically, the preset color evaluation model can analyze the color histogram in the image, calculate the proportion and position distribution of each color, and thus extract the color distribution feature.

[0058] Feature fusion refers to the process of combining multiple features into a comprehensive feature. In this step, the preset color evaluation model can fuse the contrast feature, color harmony feature, and color distribution feature to obtain a target fusion feature that contains all key color information. Among them, there are many methods of feature fusion, such as weighted average, max-min pooling, feature concatenation, etc. The specific selection can depend on actual needs. Then, the color score corresponding to the initial image can be calculated based on the target fusion feature. The color score is a quantitative index used to evaluate the color effect of the initial image, reflecting the comprehensive performance of the initial image in terms of contrast, color harmony, and color distribution. This score can be a numerical value (such as a score between 0 and 100) or a grade (such as A, B, C, etc.).

[0059] In the embodiments of the present application, by comprehensively considering multiple aspects such as the placement position of the new element, the visual center of gravity of the initial image, and the color feature, the score is calculated, so as to comprehensively evaluate the overall visual effect of the initial image and help the user to make the comprehensive effect of the generated initial image better.

[0060] In the embodiments of the present application, optionally, after obtaining the target image, the method further includes: determining the adjacent area corresponding to the new element according to the placement position of the new element in the initial image and the first preset distance range; respectively extracting the first texture feature corresponding to the new element and the second texture feature corresponding to the adjacent area, and determining the best stitching line from the adjacent area through a dynamic programming algorithm according to the first texture feature and the second texture feature; determining the fusion area of the new element in the target image according to the best stitching line and the second preset distance range, and determining the fusion weight corresponding to each pixel point according to the distance between each pixel point in the fusion area and the best stitching line; respectively calculating the first gradient information corresponding to the new element and the second gradient information corresponding to the adjacent area, and adjusting the pixel values of the pixel points in the fusion area based on the first gradient information, the second gradient information, and the fusion weights corresponding to the pixel points in the fusion area to obtain an updated target image.

[0061] In this embodiment, after integrating the newly added element into the base image to obtain the target image, the pixel values of the pixel points near the newly added element can also be processed to make the connection between the newly added element and the base image smoother and more realistic. First, an adjacent area that removes the newly added element and includes the pixels within a certain range around it can be determined according to the placement position of the newly added element and the first preset distance range. The first preset distance range can be determined according to empirical values. Then, the first texture feature is extracted from the newly added element, and the second texture feature is extracted from the adjacent area. The extraction of the texture feature can be realized by image processing algorithms (such as gray-level co-occurrence matrix, local binary pattern, etc.). After that, using the dynamic programming algorithm, according to the first texture feature and the second texture feature, the optimal stitching line between the newly added element and the base image is solved from the above determined adjacent area. This line can minimize the texture difference between the newly added element and the adjacent area, thus achieving a smooth fusion effect. Here, dynamic programming is an algorithm used to solve optimization problems.

[0062] Furthermore, an integration area that includes the optimal stitching line and the pixels within a certain range around it can be determined from the target figure according to the optimal stitching line and the second preset distance range. This area will be used for subsequent pixel value adjustment. The second preset distance range can also be determined according to empirical values. According to the distance between each pixel point in the integration area and the optimal stitching line, the integration weight corresponding to each pixel point in the integration area can be calculated. Specifically, the closer the pixel point is to the optimal stitching line, the greater the integration weight of the pixel point; the farther the pixel point is from the optimal stitching line, the smaller the integration weight of the pixel point. This integration weight is used for subsequent pixel value adjustment of the pixels in the integration area to ensure that the pixel values in the integration area can transition smoothly.

[0063] In addition, the first gradient information of the newly added elements and the second gradient information of the adjacent regions can be calculated. The gradient information generally reflects the change rate of pixel values in the image, and the above gradient information can all be calculated through image processing algorithms (such as Sobel operator, Prewitt operator, etc.). Finally, based on the first gradient information, the second gradient information, and the fusion weights corresponding to each pixel point in the fusion region, the pixel values of the pixel points in the fusion region can be adjusted. Specifically, if the difference between the first gradient value indicated by the first gradient information and the second gradient value indicated by the second gradient information is greater than a preset difference threshold, then for each pixel point in the fusion region, the new pixel value of the pixel point can be calculated by weighted averaging the pixel points within a preset range around it and the corresponding fusion weights, and the original pixel value is replaced with the new pixel value. If the difference between the first gradient value indicated by the first gradient information and the second gradient value indicated by the second gradient information is less than or equal to the preset difference threshold, then the pixel values of the pixel points in the fusion region do not need to be adjusted. Here, the preset difference threshold can be set according to requirements. It should be noted that for the pixel points at the edges of the fusion region, the pixel values do not need to be adjusted. This adjustment process can make the transition between the newly added elements and the adjacent regions smoother and more natural.

[0064] In an embodiment of the present application, optionally, after "obtaining the initial image", the method further includes: respectively calculating the first resolution corresponding to the base image and the second resolution corresponding to the newly added elements in the initial image, and based on the first resolution and the second resolution, calculating the scaling ratio corresponding to the newly added elements, determining a target scaling algorithm from the preset scaling algorithms according to the scaling ratio, and performing a scaling process on the newly added elements through the target scaling algorithm to obtain a scaled initial image; and / or, inputting the base image into a preset image style recognition model, extracting the style features corresponding to each element in the base image through the preset image style recognition model, fusing the style features corresponding to each element, determining the image style corresponding to the base image based on the fused style features, and converting the style of the newly added elements according to the image style to obtain a converted initial image.

[0065] In this embodiment, it involves the preprocessing steps before image fusion, mainly including two aspects: one is the scaling process of the newly added elements to ensure their consistency in resolution with the base image; the other is the conversion of the style of the newly added elements to make them more visually coordinated with the base image.

[0066] First, perform scaling processing on the newly added elements. First, calculate the first resolution of the base image and the second resolution of the newly added elements in the initial image. Then, based on the first resolution and the second resolution, calculate the scaling ratio of the newly added elements. According to the scaling ratio, select a suitable target scaling algorithm from the preset scaling algorithms. The preset scaling algorithms may include super-resolution reconstruction algorithms and adaptive sampling algorithms. Specifically, when magnifying the newly added elements, a super-resolution reconstruction algorithm can be used; when reducing the newly added elements, an adaptive sampling algorithm can be used. After that, use the selected target scaling algorithm to perform scaling processing on the newly added elements to obtain an initial image that matches the resolution of the base image.

[0067] Second, perform style conversion on the newly added elements. First, the base image can be input into a preset image style recognition model. This model can be a trained deep learning network that can extract style features in the image. Extract the style features corresponding to each element (or the whole) in the base image through the preset image style recognition model. Here, the style features refer to the overall style of the element, such as serious, lively, Q-version, etc. After that, fuse the extracted style features corresponding to each element. The purpose of fusion is to obtain a feature that can represent the overall style of the base image. Specifically, it can be achieved by calculating the average value, maximum value, minimum value, weighted average, etc. of the style features. Further, based on the fused style features, determine the image style corresponding to the base image. In one embodiment, the image style of the base image can be determined by comparing the fused style features with the features in the target style feature library. If there are similar style features in the target style feature library to the fused style features, the corresponding style can be determined as the style of the base image. Finally, according to the determined image style, perform style conversion on the newly added elements. After style conversion, the newly added elements can be made more visually consistent with the style of the base image. The conversion of the image style can be implemented using existing image style conversion algorithms.

[0068] The embodiments of this application perform scaling processing and style conversion on the newly added elements to ensure their consistency with the base image in terms of resolution and style, thereby ensuring the coordination between the newly added elements and the base image and improving the effect of image fusion.

[0069] In an embodiment of the present application, optionally, after "obtaining the updated target image", the method further includes: generating an image tag for the target image, marking the target image with the image tag, and storing the marked target image in a target image library, where the image tag includes at least one of the target keyword of the base image corresponding to the target image, the text-to-image prompt text, and the image description generated based on the target image; correspondingly, after "for each sub-text, performing text parsing on the sub-text to obtain a target keyword" in step 102, the method further includes: matching the target keyword with the image tags of the images stored in the target image library, determining the matching degree value corresponding to the image tag with the highest matching degree, and when the matching degree value is greater than a preset threshold, using the image corresponding to the image tag with the highest matching degree as the base image, otherwise, generating a base image corresponding to the sub-text according to the target keyword.

[0070] In this embodiment, after obtaining the smoothed target image, one or more image tags can be generated for this image. The image tags can include: the target keyword of the base image corresponding to this image, the text-to-image prompt text, and can also be the image description generated based on the target image. Specifically, the target image can be input into a large model, and the large model can be used to generate the image description of the target image.

[0071] After that, the generated image tag is used to mark the target image, and the marked target image is stored in the target image library for subsequent use or retrieval. In this way, multiple images marked with image tags can be stored in the target image library.

[0072] Therefore, before generating the base image corresponding to the sub-text each time, the target keyword of the sub-text can be matched with the image tags of the images stored in the target image library to determine the matching degree value between each image in the target image library. Here, the matching degree value can be calculated through a text similarity algorithm (such as cosine similarity, Jaccard similarity, etc.). Then, the maximum matching degree value is determined from multiple matching degree values. If this matching degree value is greater than the preset threshold, it is considered that an image highly matching the target keyword has been found, and this image is used as the base image for subsequent processing. If the matching degree value is not greater than the preset threshold, it is considered that there is no sufficiently matching image in the target image library, and at this time, a new base image can be generated according to the target keyword.

[0073] The embodiment of the present application combines image marking, storage, and image matching to implement a flexible and efficient image management system, allowing automatic retrieval of images meeting requirements according to the target keyword of the sub-text, thereby realizing the reuse of images and improving the acquisition efficiency of base images.

[0074] In an embodiment of the present application, optionally, after "determining the target image corresponding to the sub - text" in step 102, the method further includes: identifying the roles included in the sub - text and the target actions corresponding to the roles from the sub - text corresponding to the target image, and generating an animated image corresponding to the target image based on the target actions corresponding to the roles; correspondingly, "generating the picture book page corresponding to the sub - text according to each sub - text and the corresponding target image" in step 103 includes: generating the picture book page corresponding to the sub - text according to each sub - text and the animated image corresponding to the sub - text.

[0075] In this embodiment, the images in the electronic picture book can also be animated images. The animated images can be generated in the following way: First, perform text parsing on the sub - text corresponding to the target image to identify all the roles mentioned in the sub - text (subjects such as people, animals, objects, etc. that can perform actions). Role recognition can rely on natural language processing techniques, such as named entity recognition or dependency syntax analysis, to accurately extract the role information in the text. After identifying the roles, then further analyze the text to determine the target actions performed by each role. Action recognition can be obtained by parsing the verbs and their related phrases in the text to ensure accurate capture of the dynamic behavior of the roles. Based on the identified roles and the corresponding target actions, animation generation techniques (such as frame animation, vector animation, etc.) can be used to create an animated version of the target image. After generating the animated image, then create the picture book page according to each sub - text and the corresponding animated image. On the picture book page, the sub - text can be placed in an appropriate position so that readers can easily associate the text with the image when reading; the animated image can be scaled, cropped, or adjusted to meet the size and layout requirements of the picture book page. By generating picture book pages containing animated images in the embodiments of the present application, a more vivid and interesting reading experience can be provided for readers.

[0076] In an embodiment of the present application, optionally, after "obtaining the target story text" in step 101, the method further includes: identifying the dialogue text included in the target story text, determining the corresponding role for each dialogue text, and assigning different preset voices to different roles and the narration text in the target story text; for each sub-text, identifying the dialogue text, narration text included in the sub-text, and the target narration text corresponding to each dialogue text, and determining the role emotion of the role corresponding to each dialogue text based on each dialogue text and the target narration text corresponding to the dialogue text, adjusting the preset voice assigned to the role based on the role emotion, generating voice data corresponding to each dialogue text according to the preset voice after emotion adjustment, and generating voice data corresponding to the narration text according to the preset voice corresponding to the narration text; correspondingly, "generating a picture book page corresponding to the sub-text according to each sub-text and the corresponding target image" in step 103 includes: for each sub-text, generating a picture book page corresponding to the sub-text according to the sub-text, the target image corresponding to the sub-text, and the voice data.

[0077] In this embodiment, in order to further enrich the functions of the e-picture book, voice data can also be added to the e-picture book. Specifically, first, the dialogue text in the target story text can be recognized, and it can be determined which character said each dialogue text. In this way, the characters included in the entire target story text can be counted. Here, the dialogue text in the target story text can be recognized through natural language processing technology, and for each dialogue text, it can be extracted by analyzing the context of the dialogue text (such as "The little white rabbit said", "The little monkey replied", etc.). Then, a unique preset voice is assigned to each character, and at the same time, a preset voice is also assigned to the narrative text. These voices can be pre-recorded voice clips or generated through text-to-speech (TTS) technology. Among them, all the narrative texts in the target story text can be uniformly assigned a preset voice. For each sub-text, the dialogue text, narrative text, and the target narrative text corresponding to each dialogue text are recognized. Here, the target narrative text can be understood as the context text of the dialogue text. For example, the previous sentence and the next sentence of the dialogue text. The target narrative text helps to more accurately understand the context and emotion of the dialogue. For example, the dialogue text is "The little white rabbit doesn't want to play with me.", and the next sentence of this dialogue text is "The little monkey sobbed and said.", then the next sentence of the dialogue text, as the target narrative text of this dialogue text, can fully describe the character emotion of the little monkey at that time. Then, based on the dialogue text and the target narrative text, the character emotion (such as happy, sad, angry, etc.) of the character corresponding to each dialogue text is determined, and the preset voice of the character is adjusted according to this emotion to better express the emotion of the character. Specifically, it can include changing parameters such as the pitch, volume, and speech rate of the voice. Based on the preset voice after emotion adjustment, voice data corresponding to each dialogue text is generated; at the same time, voice data of the narrative text is generated according to the preset voice corresponding to the narrative text.

[0078] For each sub-text, a visually and auditorily complete picture book page is generated by combining its corresponding target image and voice data. In the embodiment of the present application, by embedding the corresponding voice data in the picture book page, so that the dialogue and narration of the characters can be heard during playback, thereby providing a more immersive reading experience for readers.

[0079] In addition, the picture book page can also be generated according to the sub-text, the voice data corresponding to the sub-text, and the animated image corresponding to the target image to simultaneously meet the multiple needs of users.

[0080] In an embodiment of the present application, optionally, after the step of "performing text segmentation on the target story text to obtain multiple sub-texts" in step 101, the method further includes: for each sub-text, according to the question matching rules corresponding to the preset question bank of each question type, performing question matching on the sub-text to obtain the target questions of the sub-text under each question type, and generating an interactive question bank corresponding to the sub-text according to the target questions, where the question types include at least one of the character learning type, concept explanation type, and thinking divergence type.

[0081] In this embodiment, in order to further enhance the educational function of the e-picture book, an interactive question bank for each sub-text can also be generated according to the content of the sub-text. Specifically, multiple question banks can be preset in advance, and each preset question bank stores questions of one question type, such as the character learning type, concept explanation type, and thinking divergence type, etc. Among them, the questions of the character learning type are designed to help children recognize and remember the characters or words in the text, such as "Please write out the synonyms of the word 'XX' in the text", "How to read the character 'X'", etc.; the questions of the concept explanation type are used to explain and elaborate on the key concepts or knowledge points in the text, such as "Please explain the meaning of the word 'XX' in the text"; the questions of the thinking divergence type are designed to stimulate students' thinking and promote the divergence and expansion of thinking, such as "Please combine the content of the text and talk about your views on the 'XX' phenomenon". In addition, a set of question matching rules are also preset for matching the sub-text with the questions in the question bank. These rules can be determined based on various technologies such as keyword matching, semantic analysis, and context understanding.

[0082] For each sub-text, question matching can be performed according to the preset question bank and question matching rules. Through question matching, one or more target questions can be generated for each sub-text under each question type. Finally, an interactive question bank is generated for each sub-text. The interactive question bank can be presented in the form of a list, table, or database, and each question can include key information such as the question text, answer, and question type. The design of the interactive question bank can promote children's learning enthusiasm. Therefore, in addition to the questions themselves, the interactive question bank can also include hints, explanations, etc. related to the questions to help children better understand the questions and answers. When generating the interactive question bank, the learning needs and interests of children can also be analyzed to perform personalized adjustment on the interactive question bank to provide learning resources and experiences that better meet the needs of students.

[0083] Further, as Figure 1 a specific implementation of the method, an embodiment of the present application provides a generating device for an e-picture book, as Figure 2 shown, the device includes:

[0084] A text segmentation module, configured to obtain first user input data, obtain a target story text according to the first user input data, and segment the target story text to obtain a plurality of sub-texts;

[0085] An image generation module, configured to, for each sub-text, perform text parsing on the sub-text to obtain target keywords, generate a basic image corresponding to the sub-text according to the target keywords, and determine a target image corresponding to the sub-text based on the basic image and second user input data, where the second user input data is used to indicate an image confirmation operation or an image modification operation;

[0086] A picture book generation module, configured to generate a picture book page corresponding to the sub-text according to each sub-text and the corresponding target image, and generate a target electronic picture book according to the picture book pages corresponding to each sub-text.

[0087] Optionally, the image generation module is configured to:

[0088] Generate a text-to-image prompt text corresponding to the sub-text according to the target keywords and a preset expansion rule, and input the text-to-image prompt text into a large model to obtain a basic image corresponding to the sub-text.

[0089] Optionally, for each sub-text in non-first sub-texts, the text-to-image prompt text further includes a guiding label constructed based on context information corresponding to the sub-text, where the context information includes the text-to-image prompt text corresponding to the previous sub-text of the current sub-text, or the target image corresponding to the previous sub-text of the current sub-text; the image generation module is further configured to:

[0090] When the context information includes the text-to-image prompt text corresponding to the previous sub-text of the current sub-text, for the first sub-text, generate a text-to-image prompt text corresponding to the first sub-text according to the target keywords corresponding to the first sub-text and a preset expansion rule, input the text-to-image prompt text corresponding to the first sub-text into a large model to obtain a basic image corresponding to the first sub-text; for each sub-text in non-first sub-texts, use the text-to-image prompt text corresponding to the previous sub-text of the current sub-text as a guiding label, generate a text-to-image prompt text corresponding to the current sub-text according to the target keywords corresponding to the current sub-text, a preset expansion rule, and the guiding label, and input the text-to-image prompt text corresponding to the current sub-text into a large model to obtain a basic image corresponding to the current sub-text;

[0091] When the context information includes the target image corresponding to the previous sub - text of the current sub - text, for the first sub - text, according to the target keyword corresponding to the first sub - text and the preset expansion rule, generate the text - to - image prompt text corresponding to the first sub - text, and input the text - to - image prompt text corresponding to the first sub - text into the large - model to obtain the basic image corresponding to the first sub - text; for each sub - text in the non - first sub - texts, use the target image corresponding to the previous sub - text of the current sub - text as the guiding label, and according to the target keyword corresponding to the current sub - text, the preset expansion rule, and the guiding label, generate the text - to - image prompt text corresponding to the current sub - text, and input the text - to - image prompt text corresponding to the current sub - text into the large - model to obtain the basic image corresponding to the current sub - text.

[0092] Optionally, when the second user input data is used to indicate an image modification operation, the image generation module is further configured to:

[0093] Based on the second user input data, determine the image modification type, where the image modification type is one of element addition, element deletion, and element modification, and the element includes at least one of a scene element, a character element, and an item element;

[0094] When the image modification type is element deletion, based on the user's area selection operation, determine the target selection area from the basic image, identify the element to be deleted included in the target selection area, and perform a deletion operation on the element to be deleted to obtain the target image corresponding to the sub - text;

[0095] When the image modification type is element modification, based on the user's area selection operation, determine the target selection area from the basic image, identify the element to be replaced included in the target selection area, perform a deletion operation on the element to be replaced to obtain the updated basic image, and integrate the new element into the basic image to obtain the target image corresponding to the sub - text;

[0096] When the image modification type is element addition, integrate the new element into the basic image to obtain the target image corresponding to the sub - text.

[0097] Optionally, the image generation module is further configured to:

[0098] Determine the placement position and scaling parameter corresponding to the new element, and according to the placement position and scaling parameter corresponding to the new element, integrate the new element into the basic image to obtain the initial image;

[0099] Perform layout analysis on the initial image to obtain a layout score, where the layout score includes at least one of the placement position score of the new element, the visual center of gravity score of the initial image, and the color score of the initial image;

[0100] When there is a score in the layout score that is less than the preset score, output a layout adjustment interaction button, and after receiving the trigger instruction of the layout adjustment interaction button, perform layout adjustment on the new element in the initial image, and use the adjusted initial image as the target image;

[0101] When all scores in the layout score are greater than or equal to the preset score, use the initial image as the target image.

[0102] Optionally, the image generation module is further configured to:

[0103] Based on the placement position of the new element in the initial image, determine the target analysis area corresponding to the placement position, identify the original elements included in the target analysis area, determine the element intersection between the new element and the original elements according to the placement position and the target positions corresponding to the original elements, calculate the area ratio of the element intersection to the new element, and calculate the placement position score corresponding to the new element based on the area ratio; and / or,

[0104] Based on the preset composition principle and / or user image preference, determine the ideal center of gravity corresponding to the initial image, and according to the position, element size, and element color corresponding to each element in the initial image, determine the element weight corresponding to each element. According to the element center point position and element weight of each element, calculate the weighted position corresponding to the element, add up the weighted positions corresponding to each element to obtain the visual center of gravity corresponding to the initial image, and calculate the visual center of gravity score corresponding to the initial image based on the ideal center of gravity and the visual center of gravity; and / or,

[0105] Input the initial image into a preset color evaluation model, respectively extract the contrast feature, color harmony feature, and color distribution feature of the initial image through the preset color evaluation model, perform feature fusion on the contrast feature, color harmony feature, and color distribution feature to obtain a target fusion feature, and calculate the color score corresponding to the initial image according to the target fusion feature.

[0106] Optionally, the device further includes an image update module, and the image update module is configured to:

[0107] After obtaining the target image, determine the adjacent area corresponding to the new element according to the placement position of the new element in the initial image and the first preset distance range;

[0108] Extract the first texture feature corresponding to the newly added element and the second texture feature corresponding to the adjacent region respectively. According to the first texture feature and the second texture feature, determine the optimal stitching line from the adjacent region through the dynamic programming algorithm.

[0109] Determine the fusion region of the newly added element in the target image according to the optimal stitching line and the second preset distance range, and determine the fusion weight corresponding to each pixel point according to the distance between each pixel point in the fusion region and the optimal stitching line.

[0110] Calculate the first gradient information corresponding to the newly added element and the second gradient information corresponding to the adjacent region respectively. Based on the first gradient information, the second gradient information and the fusion weight corresponding to each pixel point in the fusion region, adjust the pixel values of the pixel points in the fusion region to obtain the updated target image.

[0111] Optionally, the device further includes an initial image processing module, and the initial image processing module is used for:

[0112] After obtaining the initial image, calculate the first resolution corresponding to the base image and the second resolution corresponding to the newly added element in the initial image respectively, and calculate the scaling ratio corresponding to the newly added element based on the first resolution and the second resolution. According to the scaling ratio, determine the target scaling algorithm from the preset scaling algorithms, and perform scaling processing on the newly added element through the target scaling algorithm to obtain the scaled initial image; and / or,

[0113] Input the base image into a preset image style recognition model, extract the style features corresponding to each element in the base image through the preset image style recognition model, and fuse the style features corresponding to each element. Based on the fused style features, determine the image style corresponding to the base image, and convert the style of the newly added element according to the image style to obtain the converted initial image.

[0114] Optionally, the device further includes:

[0115] An image storage module, which is used for generating an image label for the target image after obtaining the updated target image, marking the target image through the image label, and storing the marked target image in a target image library, where the image label includes at least one of the target keyword of the base image corresponding to the target image, the text-to-image prompt text, and the image description generated based on the target image;

[0116] Correspondingly, the device further includes:

[0117] An image matching module, for each sub - text, after parsing the sub - text to obtain target keywords, matching the target keywords with the image tags of the images stored in the target image library, determining the matching degree value corresponding to the image tag with the highest matching degree, and when the matching degree value is greater than a preset threshold, using the image corresponding to the image tag with the highest matching degree as the base image; otherwise, generating a base image corresponding to the sub - text according to the target keywords.

[0118] Optionally, the device further includes:

[0119] An animated image generation module, after determining the target image corresponding to the sub - text, identifying the characters included in the sub - text and the target actions corresponding to the characters from the sub - text corresponding to the target image, and generating an animated image corresponding to the target image based on the target actions corresponding to the characters.

[0120] Correspondingly, the picture book generation module is used for:

[0121] Generating a picture book page corresponding to each sub - text according to each sub - text and the animated image corresponding to the sub - text.

[0122] Optionally, the device further includes a voice generation module, and the voice generation module is used for:

[0123] After obtaining the target story text, identifying the dialogue text included in the target story text, determining the characters corresponding to each dialogue text, and assigning different preset voices to different characters and the narration text in the target story text.

[0124] For each sub - text, identifying the dialogue text, narration text included in the sub - text, and the target narration text corresponding to each dialogue text, determining the character emotions of the characters corresponding to each dialogue text according to each dialogue text and the target narration text corresponding to the dialogue text, adjusting the emotions of the preset voices assigned to the characters based on the character emotions, generating voice data corresponding to each dialogue text according to the preset voices after emotion adjustment, and generating voice data corresponding to the narration text according to the preset voice corresponding to the narration text.

[0125] Correspondingly, the picture book generation module is further used for:

[0126] For each sub - text, generating a picture book page corresponding to the sub - text according to the sub - text, the target image corresponding to the sub - text, and the voice data.

[0127] Optionally, the device further includes:

[0128] An interactive question bank generation module is used to perform text segmentation on the target story text to obtain multiple sub-texts. Then, for each sub-text, according to the question matching rules corresponding to the preset question banks of each question type, the sub-text is matched with questions to obtain the target questions of the sub-text under each question type, and an interactive question bank corresponding to the sub-text is generated, where the question type includes at least one of the character learning type, the concept explanation type, and the thinking divergence type.

[0129] It should be noted that for other corresponding descriptions of each functional unit involved in the electronic picture book generation device provided in the embodiments of the present application, reference can be made to Figure 1 the corresponding descriptions in the method, which will not be elaborated here.

[0130] The embodiments of the present application further provide a computer device, which can specifically be a personal computer, a server, a network device, etc. As Figure 3 shown, the computer device includes a bus, a processor, a memory, and a communication interface, and may further include an input / output interface and a display device. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store location information. The network interface of the computer device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements the steps in the method embodiments.

[0131] Those skilled in the art can understand that Figure 3 the structure shown in

[0132] is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0133] In one embodiment, a computer-readable storage medium is provided. The computer-readable storage medium may be non-volatile or volatile, and has a computer program stored thereon. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0134] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0135] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0136] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0137] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A method for generating an electronic picture book, characterized in that: include: Acquire first user input data, obtain a target story text according to the first user input data, and perform text segmentation on the target story text to obtain a plurality of sub-texts; For each sub-text, performing text parsing on the sub-text to obtain target keywords, generating a basic image corresponding to the sub-text according to the target keywords, and determining a target image corresponding to the sub-text based on the basic image and second user input data, wherein the second user input data is used to indicate an image confirmation operation or an image modification operation; According to each sub-text and the corresponding target image, a picture book page corresponding to the sub-text is generated, and according to the picture book page corresponding to each sub-text, a target electronic picture book is generated.

2. The method according to claim 1, characterized in that The step of generating a basic image corresponding to the subtext according to the target keyword includes: According to the target keyword and the preset expansion rule, a text image prompt text corresponding to the sub-text is generated, and the text image prompt text is input into the large model to obtain a basic image corresponding to the sub-text.

3. The method according to claim 2, characterized in that For each sub-text other than the first sub-text, the text image prompt text further includes a guide label constructed based on context information corresponding to the sub-text, wherein the context information includes a text image prompt text corresponding to a previous sub-text of the current sub-text, or a target image corresponding to a previous sub-text of the current sub-text; the text image prompt text corresponding to the sub-text is generated according to the target keyword and the preset expansion rule, and the text image prompt text is input into the large model to obtain the basic image corresponding to the sub-text, including: When the context information includes a text image prompt text corresponding to a previous subtext of the current subtext, for the first subtext, according to the target keyword corresponding to the first subtext and the preset expansion rule, the text image prompt text corresponding to the first subtext is generated, and the text image prompt text corresponding to the first subtext is input into the large model to obtain the basic image corresponding to the first subtext; for each subtext other than the first subtext, the text image prompt text corresponding to the previous subtext of the current subtext is used as a guide label, according to the target keyword corresponding to the current subtext, the preset expansion rule and the guide label, the text image prompt text corresponding to the current subtext is generated, and the text image prompt text corresponding to the current subtext is input into the large model to obtain the basic image corresponding to the current subtext; When the context information includes a target image corresponding to a previous subtext of a current subtext, for the first subtext, a text image prompt text corresponding to the first subtext is generated according to the target keyword corresponding to the first subtext and a preset expansion rule, and the text image prompt text corresponding to the first subtext is input into the large model to obtain a basic image corresponding to the first subtext; for each subtext other than the first subtext, the target image corresponding to the previous subtext of the current subtext is used as a guide label, and a text image prompt text corresponding to the current subtext is generated according to the target keyword corresponding to the current subtext, the preset expansion rule and the guide label, and the text image prompt text corresponding to the current subtext is input into the large model to obtain a basic image corresponding to the current subtext.

4. The method according to any one of claims 1 to 3, characterized in that When the second user input data is used to indicate an image modification operation, determining a target image corresponding to the subtext based on the base image and the second user input data includes: Determine an image modification type based on the second user input data, wherein the image modification type is one of element addition, element deletion, and element modification, and the element includes at least one of a scene element, a character element, and an object element; When the image modification type is element deletion, based on the user's area selection operation, a target selection area is determined from the basic image, and the elements to be deleted contained in the target selection area are identified, and a deletion operation is performed on the elements to be deleted to obtain a target image corresponding to the subtext; When the image modification type is element modification, based on the user's area selection operation, a target selection area is determined from the basic image, and the elements to be replaced contained in the target selection area are identified, a deletion operation is performed on the elements to be replaced to obtain an updated basic image, and the newly added elements are integrated into the basic image to obtain a target image corresponding to the subtext; When the image modification type is element addition, the newly added element is integrated into the basic image to obtain a target image corresponding to the subtext.

5. The method according to claim 4, characterized in that The step of integrating the newly added elements into the basic image to obtain a target image corresponding to the subtext includes: Determine a placement position and a scaling parameter corresponding to the newly added element, and integrate the newly added element into the basic image according to the placement position and the scaling parameter corresponding to the newly added element to obtain an initial image; Performing a layout analysis on the initial image to obtain a layout score, wherein the layout score includes at least one of a placement position score of the newly added element, a visual center of gravity score of the initial image, and a color score of the initial image; When there is a score less than a preset score in the layout scores, a layout adjustment interaction button is output, and after receiving a trigger instruction of the layout adjustment interaction button, the layout of the newly added elements in the initial image is adjusted, and the adjusted initial image is used as the target image; When all scores in the layout score are greater than or equal to a preset score, the initial image is used as a target image.

6. The method according to claim 5, characterized in that The performing layout analysis on the initial image to obtain a layout score includes: Based on the placement position of the newly added element in the initial image, a target analysis area corresponding to the placement position is determined, and original elements included in the target analysis area are identified, and according to the placement position and the target position corresponding to the original element, an element intersection between the newly added element and the original element is determined, and an area ratio of the element intersection to the newly added element is calculated, and a placement position score corresponding to the newly added element is calculated based on the area ratio; and / or, Based on preset composition principles and / or user image preferences, determine the ideal center of gravity corresponding to the initial image, and determine the element weight corresponding to each element according to the position, element size and element color corresponding to each element in the initial image, calculate the weighted position corresponding to the element according to the element center point position and element weight of each element, add the weighted positions corresponding to each element to obtain the visual center of gravity corresponding to the initial image, and calculate the visual center of gravity score corresponding to the initial image based on the ideal center of gravity and the visual center of gravity; and / or, The initial image is input into a preset color evaluation model, and the contrast features, color harmony features, and color distribution features of the initial image are respectively extracted through the preset color evaluation model. The contrast features, color harmony features, and color distribution features are fused to obtain target fusion features, and the color score corresponding to the initial image is calculated based on the target fusion features.

7. The method according to claim 5, characterized in that After obtaining the target image, the method further includes: Determining an adjacent area corresponding to the newly added element according to a placement position of the newly added element in the initial image and a first preset distance range; Extracting first texture features corresponding to the newly added elements and second texture features corresponding to the adjacent areas respectively, and determining an optimal splicing line from the adjacent areas through a dynamic programming algorithm according to the first texture features and the second texture features; Determining a fusion region of the newly added element in the target image according to the optimal stitching line and a second preset distance range, and determining a fusion weight corresponding to each pixel point in the fusion region according to a distance between each pixel point and the optimal stitching line; The first gradient information corresponding to the newly added element and the second gradient information corresponding to the adjacent area are calculated respectively. Based on the first gradient information, the second gradient information and the fusion weights corresponding to each pixel in the fusion area, the pixel values ​​of the pixel points in the fusion area are adjusted to obtain an updated target image.

8. The method according to claim 5, characterized in that After obtaining the initial image, the method further comprises: respectively calculating a first resolution corresponding to the basic image and a second resolution corresponding to the newly added element in the initial image, and calculating a scaling ratio corresponding to the newly added element based on the first resolution and the second resolution, determining a target scaling algorithm from a preset scaling algorithm according to the scaling ratio, and scaling the newly added element using the target scaling algorithm to obtain a scaled initial image; and / or, The basic image is input into a preset image style recognition model, the style features corresponding to each element in the basic image are extracted through the preset image style recognition model, and the style features corresponding to the elements are fused, and the image style corresponding to the basic image is determined based on the fused style features, and the style of the newly added elements is converted according to the image style to obtain the converted initial image.

9. The method according to claim 7, characterized in that: After obtaining the updated target image, the method further includes: Generate an image tag for the target image, mark the target image by the image tag, and store the marked target image in a target image library, wherein the image tag includes at least one of a target keyword of a basic image corresponding to the target image, a text image prompt text, and an image description generated based on the target image; Correspondingly, for each sub-text, after parsing the sub-text to obtain the target keyword, the method further includes: The target keyword is matched with the image tags of the images stored in the target image library, and the matching value corresponding to the image tag with the highest matching degree is determined. When the matching value is greater than a preset threshold, the image corresponding to the image tag with the highest matching degree is used as the basic image. Otherwise, the basic image corresponding to the sub-text is generated according to the target keyword.

10. The method according to claim 1, characterized in that After determining the target image corresponding to the subtext, the method further includes: From the subtext corresponding to the target image, identifying the character contained in the subtext and the target action corresponding to the character, and generating an animation image corresponding to the target image based on the target action corresponding to the character; Accordingly, generating a picture book page corresponding to each sub-text according to the corresponding target image includes: According to each sub-text and the animated image corresponding to the sub-text, a picture book page corresponding to the sub-text is generated.

11. The method according to claim 1, characterized in that: After obtaining the target story text, the method further includes: Identify the dialogue texts contained in the target story text, determine the characters corresponding to each dialogue text, and assign different preset voices to different characters and the narration texts in the target story text; For each sub-text, identifying the dialogue text, narration text, and target narration text corresponding to each dialogue text contained in the sub-text, and determining the role emotion of the character corresponding to each dialogue text according to each dialogue text and the target narration text corresponding to the dialogue text, performing emotion adjustment on the preset sound assigned to the character based on the role emotion, generating voice data corresponding to each dialogue text according to the preset sound after the emotion adjustment, and generating voice data corresponding to the narration text according to the preset sound corresponding to the narration text; Accordingly, generating a picture book page corresponding to each sub-text according to the corresponding target image includes: For each sub-text, a picture book page corresponding to the sub-text is generated according to the sub-text, a target image corresponding to the sub-text, and voice data.

12. The method according to claim 1, characterized in that After segmenting the target story text to obtain a plurality of sub-texts, the method further includes: For each sub-text, question matching is performed on the sub-text according to the question matching rules corresponding to the preset question bank of each question type to obtain the target question of the sub-text under each question type, and based on the target question, an interactive question bank corresponding to the sub-text is generated, wherein the question type includes at least one of a new word learning type, a concept explanation type and a divergent thinking type.

13. A device for generating an electronic picture book, characterized in that: include: A text segmentation module, used to obtain first user input data, obtain a target story text according to the first user input data, and perform text segmentation on the target story text to obtain a plurality of sub-texts; an image generation module, configured to perform text parsing on each subtext to obtain a target keyword, generate a basic image corresponding to the subtext according to the target keyword, and determine a target image corresponding to the subtext based on the basic image and second user input data, wherein the second user input data is used to indicate an image confirmation operation or an image modification operation; The picture book generation module is used to generate a picture book page corresponding to each sub-text according to the sub-text and the corresponding target image, and to generate a target electronic picture book according to the picture book page corresponding to each sub-text.

14. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.

15. A computer device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 12 is implemented.