A continuous children's picture book image generation method based on dynamic memory enhancement

By decoupling role identity and posture features through frequency domain high-pass filtering and dual-source dynamic memory, and combining multimodal visual correction mechanism, the problems of noise interference, feature coupling and detail loss in the generation of long sequence children's picture books are solved, and high-fidelity and coherent picture book image generation is achieved.

CN122636775APending Publication Date: 2026-08-25HANGZHOU NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610753230.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-28
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing technologies, when applied to long-sequence children's picture book generation scenarios, suffer from noise interference in training data, easy coupling between character identity features and spatial posture, discontinuity and loss of details in the temporal features of long-sequence generation, and lack of automated feedback on fine-grained structural rationality, resulting in discontinuous and visually distorted picture book images.

Method used

Frequency domain high-pass filtering and dual-source dynamic memory are introduced to decouple role identity features and historical action postures. A multimodal visual automatic correction mechanism is used to explicitly punish physiological structural distortions. Cross-image feature fusion and temporal decay bias matrix are used to achieve high-fidelity texture transfer and environmental state smoothing.

Benefits of technology

It improves the narrative coherence, aesthetic unity, and visual rationality of long-sequence picture book generation, and enhances the practical usability of the generated images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122636775A_ABST
    Figure CN122636775A_ABST
Patent Text Reader

Abstract

The application discloses a continuous children's picture book image generation method based on dynamic memory enhancement. The application firstly performs multi-stage cascading cleaning on original children's picture book images to construct a high-quality image-text pair dataset; inserts a consistency enhancement low-rank adapter into a pre-trained text conditional image generation base model to construct a dual-source dynamic feature memory bank; performs frequency domain decoupling processing on latent feature maps corresponding to reference images and historical images, extracts local edge, texture details and local structure information, and combines a time sequence attenuation bias to perform memory enhancement attention fusion; constructs an automatic preference sample based on a multi-modal visual evaluation model, trains a preference alignment low-rank adapter, and obtains a continuous picture book image generation model based on preference alignment. The application can improve the role appearance consistency, page coherence and visual structure rationality in the continuous picture book image generation process, and reduce the occurrence probability of role drift, style mutation and local structure anomaly in long sequence generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence and computer vision technology, and relates to a method for generating continuous children's picture book images based on dynamic memory enhancement, which is a continuous image generation method based on text conditional image generation model. Background Technology

[0002] In recent years, text-conditional image generation technology based on text-conditional image generation models has made significant progress. However, when dealing with continuous image generation scenarios such as children's picture books that require long narrative sequences, high character consistency, and complex visual logic, existing technologies still have the following four significant limitations:

[0003] First, the training data suffers from deep-seated noise interference. The open-source picture book images acquired in practice typically contain non-narrative pages and large areas of text obscuring the text. Conventional preprocessing methods struggle to automatically repair background textures without compromising the visual structure, making the model's feature learning susceptible to interference from irrelevant text and layout, severely impacting the narrative coherence of subsequent generation.

[0004] Second, character identity features are highly coupled with spatial pose. Existing topic-driven fine-tuning methods based on reconstruction loss (such as the DreamBooth scheme, see: Ruiz, Nataniel, et al. “Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.” Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2023.) or cross-image attention mechanisms that directly aggregate all features often confuse high-frequency features of the character (such as facial details and clothing textures) with low-frequency features (such as large color blocks in the overall background, limb contours, and action poses) in the spatial domain when extracting features from the reference image. This causes the target image to be easily "forcibly pulled" by the old actions of the reference image when the model transfers features across images, making the character identity strongly bound to its specific pose or background in the training set. When generating new and complex actions, the model's zero-shot generalization ability drops sharply, and character pose stiffness or command failure are very likely to occur.

[0005] Third, long sequence generation suffers from temporal feature discontinuities and detail loss. Existing feature injection mechanisms (such as ControlNet-based methods, see: Zhang, Lvmin, Anyi Rao, and Maneesh Agrawala. "Adding conditional control to text-to-image diffusion models." Proceedings of the IEEE / CVF international conference on computer vision. 2023.) often focus on extracting high-level contour semantics from a single reference image, making it difficult to accurately replicate low-level high-frequency textures and global color distribution. Furthermore, because existing methods lack long-term dynamic memory and smooth temporal state evolution mechanisms for long sequence generation, the model can only perform isolated "image-to-image" single-point references when continuously generating multiple frames. Micro-environmental states (such as the gradual change of light and shadow over time, and the gradual wear and tear of clothing in the plot) cannot be naturally transmitted, easily leading to unnatural abrupt changes in tone and texture, severely damaging the aesthetic unity and temporal coherence of long sequence images.

[0006] Fourth, there is a lack of a fine-grained automated feedback mechanism for structural rationality. The denoising process of text-conditional image generation models focuses on fitting the statistical distribution of local pixels, lacking an understanding of the rationality of local structures. Relying on manually constructing static positive and negative sample pairs datasets is not only costly to build and difficult to cover long-tailed distortion scenarios, but also makes it difficult for conventional supervised fine-tuning to explicitly discriminate and penalize high-confidence errors that violate real physiological structures, such as polydactyly and reverse joint bending, resulting in generated images that are highly susceptible to structural visual distortions.

[0007] In summary, existing technologies face shortcomings in continuous picture book generation scenarios, such as high data noise, easy coupling of frequency domain features, easy loss of texture, and lack of automated correction. There is an urgent need in this field for a new generation method to achieve high-fidelity, high-coherence, and consistent generation of characters in continuous picture book images that conform to human visual cognitive logic. Summary of the Invention

[0008] This invention aims to address the shortcomings of existing text-based conditional image generation models in generating long sequences of children's picture books. These shortcomings include deep noise interference in training data, high coupling between character identity features and specific spatial postures, temporal feature discontinuities and loss of details leading to inconsistencies in detail and style, and a lack of fine-grained character structural rationality and physiological logic constraints that easily cause structural visual distortions. The invention provides a continuous children's picture book image generation method based on dynamic memory enhancement. This invention focuses on introducing frequency domain high-pass filtering and a dual-source long-term dynamic memory library to decouple character identity features from historical action postures, achieving high-fidelity unidirectional texture transfer across images and smooth transitions in environmental states. Furthermore, it innovatively introduces multimodal visual automatic correction, using a direct preference optimization algorithm to explicitly penalize and suppress physiological structural distortions in the underlying generation logic, thereby significantly improving the narrative coherence, aesthetic unity, and automated visual rationality of long sequence picture book generation.

[0009] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0010] Step (1) Collect raw image data and preprocess it to obtain image-text pair dataset. The raw image data is children's picture book image data.

[0011] Step (2) Construct a text conditional image generation model with cross-image feature fusion capability;

[0012] (2-1) Insert a consistency enhancement low-rank adapter into the pre-trained text conditional image generation base model, keep the backbone parameters of the pre-trained text conditional image generation base model frozen, and update only the parameters of the consistency enhancement low-rank adapter during the training phase.

[0013] (2-2) Constructing latent feature maps for multi-source images:

[0014] The image-text pair dataset is organized into training sample groups according to role association and page order. Each training sample group includes a target image-text pair, one or more reference image-text pairs, and multiple historical image-text pairs. The target image, reference image, and historical image in the target image-text pair, reference image-text pair, and historical image-text pair are encoded using a pre-trained variational autoencoder to map the target image, reference image, and historical image to the latent space to obtain the corresponding latent feature maps.

[0015] (2-3) Construct a dual-source dynamic feature memory:

[0016] The dual-source dynamic feature memory includes a core anchor feature storage area and a historical evolution feature storage area; wherein, the core anchor feature storage area is used to store the latent feature map corresponding to the reference image, and the historical evolution feature storage area is used to store the latent feature map corresponding to the historical image, and the latent feature map of the historical image corresponding to the most recent historical page is retained according to a preset window length.

[0017] (2-4) Construct an extended feature pool based on frequency domain decoupling:

[0018] Frequency domain decoupling is performed on the latent feature maps of reference images and historical images in the dual-source dynamic feature memory to extract high-frequency features that characterize local edges, texture details, and local structural information. The high-frequency features are then segmented and linearly projected to obtain high-frequency reference visual feature sequences and high-frequency historical evolution feature sequences. These sequences are then input into the key projection layer and value projection layer of the pre-trained text conditional image generation base model to obtain reference image key-value features and historical image key-value features, respectively.

[0019] The latent feature map of the target image is processed by block segmentation and linear projection to obtain the target visual feature sequence. The target visual feature sequence is then input into the key projection layer and value projection layer of the pre-trained text conditional image generation base model to obtain the key-value features of the target image itself.

[0020] The key features of the target image itself, the key features of the reference image, and the key features of the historical images are concatenated in the sequence dimension to construct an extended feature pool;

[0021] (2-5) Memory enhancement and attention fusion based on temporal decay bias:

[0022] The target visual feature sequence is input into the query projection layer of the pre-trained text conditional image generation base model to obtain the target image query features; a temporal decay bias is constructed based on the page distance between the historical image and the target image, and the historical image features that are farther away from the target image have a lower impact on the attention calculation; memory-enhanced attention fusion is performed based on the target image query features, the key-value features in the extended feature pool, and the temporal decay bias to obtain the memory-enhanced target image generation features.

[0023] (2-6) Parameter optimization based on target image reconstruction loss:

[0024] The target image is trained by denoising and reconstructing the generated features of the memory-enhanced target image. The reconstruction loss is used to update the parameters of the consistency-enhanced low-rank adapter. After training, the pre-trained text conditional image generation base model, which superimposes the consistency-enhanced low-rank adapter, the dual-source dynamic feature memory, the extended feature pool, and the memory-enhanced attention fusion module, becomes a text conditional image generation model with cross-image feature fusion capability.

[0025] Step (3) Construct a continuous picture book image generation model based on preference alignment;

[0026] (3-1) Generate a candidate image set:

[0027] For the current picture book page, first obtain the text prompts for the current page. and reference image information Read the current state of the dual-source dynamic feature memory. Construct the composite generation conditions for the current page , By changing random noise, sampling parameters, or repeatedly performing the sampling process, multiple candidate image samples are generated to obtain a candidate image set.

[0028] (3-2) Construct an automated preference sample dataset based on a multimodal visual evaluation model:

[0029] The candidate image set is input into the multimodal visual evaluation model. Combined with the preset visual evaluation instructions, the model evaluates the visual structure rationality, role consistency and page coherence of the samples in the candidate image set and obtains the corresponding preference evaluation score. Based on the preference evaluation score, positive and negative samples are selected from the candidate image samples under the same composite generation conditions. The composite generation conditions, positive samples and negative samples are combined to construct preference triplets to form an automated preference sample dataset.

[0030] (3-3) The text conditional image generation model is used as the reference model, and the parameters of the pre-trained text conditional image generation base model and the consistency enhancement low-rank adapter are frozen; a trainable preference alignment low-rank adapter is superimposed on the reference model as a policy model, and in the preference optimization stage, only the parameters of the preference alignment low-rank adapter are updated.

[0031] (3-4) The policy model is directly optimized based on the automated preference sample dataset, so that the policy model improves the preference score of positive samples and reduces the preference score of negative samples compared with the reference model. After training, a continuous picture book image generation model containing consistency enhancement low-rank adapter and preference alignment low-rank adapter is obtained.

[0032] Step (4) Generate continuous children's picture book images;

[0033] (4-1) Obtain the page text sequence and reference image information of children's picture books; the page text sequence includes multiple page text prompts, each of which is used to describe the character's actions, scene content and picture style in the corresponding picture book page; the reference image information is used to provide the target character's identity information, appearance details and drawing style information;

[0034] (4-2) Initialize the dual-source dynamic feature memory based on the reference image information and write the reference latent feature map corresponding to the reference image into the core anchor point feature storage area;

[0035] (4-3) For the current picture book page to be generated, read the text prompts, reference image information and the current state of the dual-source dynamic feature memory bank, and construct composite generation conditions; input the composite generation conditions into the continuous picture book image generation model to generate the current page image;

[0036] (4-4) Encode the current page image to obtain the current page potential feature map and write it into the historical evolution feature storage area; repeat the current page image generation and historical evolution feature storage area update process according to the page order until a complete continuous picture book image sequence is obtained.

[0037] The beneficial effects of this invention include: by proposing a cascaded multi-level data cleaning pipeline that includes semantic usability filtering, optical character localization, and background redrawing and restoration, the background is automatically extracted and reconstructed without damaging the underlying physical structure of the image, eliminating noise pollution at the data source and greatly improving the model's ability to extract specific artistic styles and pure features; through a cross-image unidirectional feature collaboration mechanism based on frequency domain decoupling and a dual-source dynamic memory, frequency domain high-pass filtering is introduced during feature extraction, effectively filtering out low-frequency pose interference representing limb contours. Simultaneously, a temporal decay bias matrix is ​​introduced in the temporal scheduling, making the model highly dependent on recent historical frames for micro-environmental states, while relying on the initial reference image for core facial textures, achieving deep decoupling between high-fidelity character identity and spatial action pose, solving the problem of the inability to continuously and smoothly transmit micro-states in long sequence generation; finally, by introducing a multimodal proxy automated preference alignment strategy, combined with a direct preference optimization algorithm, explicit logical judgment and punishment are applied to high-confidence errors that violate real physiological structures, helping to suppress structural visual distortion and improving the practical application usability and spatial logical rationality of continuous multi-image generation. Attached Figure Description

[0038] Figure 1 This is a schematic diagram of the overall process of a continuous children's picture book image generation method based on dynamic memory enhancement according to the present invention. Detailed Implementation

[0039] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0040] A method for generating continuous children's picture book images based on dynamic memory enhancement, such as Figure 1 As shown, it includes the following steps:

[0041] Step (1) Collect raw image data and preprocess it to obtain image-text pair dataset. The raw image data is children's picture book image data.

[0042] (1-1) Collect raw image data to obtain raw picture book image data covering a variety of art styles and character types.

[0043] This embodiment utilizes web crawling technology or public dataset interfaces to acquire children's picture book image data covering various art styles (e.g., watercolor, oil painting, cartoon, etc.) and character types (e.g., anthropomorphic animals, fantasy creatures, etc.) from open-source digital libraries and picture book resource databases. In this embodiment, approximately 10,000 picture books in the public domain (e.g., CC0 or CC-BY licenses) are preferably collected to ensure the style generalization ability of the pre-trained text-conditional image generation base model (e.g., the FLUX.1 model).

[0044] (1-2) Calculate the narrative confidence of each original image data, remove non-narrative pages with confidence below a preset threshold, and obtain usable image data.

[0045] Given that the original data contains a large amount of non-narrative noise such as covers, tables of contents, and blank pages, this embodiment constructs and fine-tunes a dedicated binary classifier based on a pre-trained image-text matching model (such as the CLIP model). In the automated screening phase, utilizing... Calculate each input image Narrative confidence And set a preset threshold. (In this embodiment, 0.9 is preferred). The system only retains values ​​that satisfy... Images are used to precisely remove non-narrative pages that do not contain the main characters or core scenes of the story.

[0046] (1-3) Text removal: The optical character detection model is used to locate the text region in the available image data, generate a binary mask, input the available image data and the binary mask into the image redrawing model, and use global context information to reconstruct the background texture obscured by the text, and output a clean image.

[0047] To address the common problem of text obscuring the image in picture book pages, this embodiment employs a combined strategy of "optical character detection + background redrawing and restoration." First, an optical character detection model (such as the CRAFT model) is used to accurately locate the pixel-level coordinates of all text regions in the image, generating a binary mask. (in Represents a text area. (representing non-text areas); then, the original image is compared with a binary mask. The image is synchronously input into a large receptive field image redrawing model (such as the LaMa model). By utilizing the model's understanding of global context information, the high-frequency background textures that are obscured by text are reconstructed and filled in, and a clean image without text interference is output.

[0048] (1-4) Detect the pixel resolution and aspect ratio of the clean image, and remove distorted images with resolution below the preset size limit or aspect ratio outside the preset range to obtain a high-quality image.

[0049] The redrawn clean image is scanned at physical resolution, and image height is automatically removed. Or width Low-resolution images smaller than a preset size limit (e.g., 512 pixels) are used; aspect ratio verification is performed simultaneously to remove distorted images with extremely unbalanced aspect ratios, preventing geometric deformation interference during the feature extraction stage of the generative network.

[0050] (1-5) Calculate the Hamming distance of structured features between any two high-quality images using the perceptual hash algorithm, remove approximately duplicate images with similarity less than the similarity threshold, and obtain the final image data.

[0051] To avoid overfitting the model due to near-redundant data, content-aware deduplication is performed on clean images. This embodiment uses a perceptual hashing algorithm to extract the structured feature hash fingerprint of each image and calculates the Hamming distance between any two images. Set a Hamming distance similarity threshold (preferably 10 in this embodiment). When When the two are determined to be highly similar in visual semantics, one is retained and the duplicate copy is removed.

[0052] (1-6) Use a multimodal large language model to perform visual understanding on the final image data, generate refined text descriptions containing character appearance, actions and scene details, and obtain image-text pair datasets.

[0053] We utilize multimodal large language models (such as the Florence-2-large model) to perform deep visual semantic analysis on the deduplicated clean images, automatically generating fine-grained natural language text descriptions that include character appearance features, current action semantics, and scene spatial layout. We then bind each generated text description to a corresponding clean image to construct a standardized dataset of picture book narrative image-text pairs.

[0054] Step (2) Construct a text conditional image generation model with cross-image feature fusion capability.

[0055] (2-1) Insert a consistency enhancement low-rank adapter into the pre-trained text conditional image generation pedestal model, keep the backbone parameters of the pre-trained text conditional image generation pedestal model frozen, and update only the parameters of the consistency enhancement low-rank adapter during the training phase.

[0056] The consistency enhancement low-rank adapter is used to learn the cross-image feature mapping relationships and feature interaction relationships between reference images, historical images, and target images. This allows the target image to incorporate character identity information, clothing texture information, and painting style information from the reference image during generation, and combines local state evolution information from historical images to improve the consistency of character appearance and page coherence in continuous picture book image generation. The consistency enhancement low-rank adapter is a parameter-efficient fine-tuning module, comprising at least one set of low-rank parameter matrices and corresponding scaling coefficients. This adapter is inserted in parallel into the bypass of the original linear transformation layer of the pre-trained text conditional image generation pedestal model. While keeping the original weight matrix frozen, it generates trainable weight increments through the low-rank parameter matrices to supplement and correct the feature transformation results of the pre-trained text conditional image generation pedestal model. Specifically, for the original weight matrix in the pre-trained text conditional image generation pedestal model, the consistency enhancement low-rank adapter constructs weight increments through the product of two low-rank matrices and applies these weight increments to the input features, thus forming a superimposed output of the original feature transformation results and the low-rank increment features.

[0057] Keep the backbone parameters of the pre-trained text conditional image generation pedestal model frozen, and insert consistency-enhancing low-rank adapters into the attention projection layer and feedforward network module of the pre-trained text conditional image generation pedestal model. The attention projection layer includes a query projection layer, a key projection layer, a value projection layer, and an output projection layer; the feedforward network module includes a multilayer perceptron module. During the training phase, only the consistency-enhanced low-rank adapter is updated. The parameters of the adapter are preferably set in this embodiment. The rank is 16, the scaling factor is 16, and the parameters are initialized using a normal distribution.

[0058] (2-2) Constructing latent feature maps for multi-source images:

[0059] The image-text pair dataset is organized into training sample groups according to role association and page order. Each training sample group includes one target image-text pair, one or more reference image-text pairs, and multiple historical image-text pairs. The target image-text pair includes the target image and its corresponding text description, providing information about the target sample to be generated. The reference image-text pair includes the reference image and its corresponding text description, providing character identity and painting style information. The historical image-text pair includes historical images and their corresponding text descriptions, providing state evolution information before the target page. State evolution information includes at least one of character pose changes, local appearance changes, and scene changes. Historical images are those preceding the target image. Page (Preferred in this embodiment) 3) Continuous page images; when there is no continuous page relationship in the training data, historical images can also be composed of associated images with the same role identity as the target image.

[0060] The target image, reference image, and historical image in the target image-text pair, reference image-text pair, and historical image-text pair are encoded using a pre-trained variational autoencoder (e.g., VAE) to map the target image, reference image, and historical image to the latent space, thereby obtaining the latent feature maps of the target image, reference image, and historical image.

[0061] (2-3) Construct a dual-source dynamic feature memory:

[0062] The dual-source dynamic feature memory includes a core anchor feature storage area and a historical evolution feature storage area; wherein, the core anchor feature storage area is used to store the latent feature map corresponding to the reference image, and the historical evolution feature storage area is used to store the latent feature map corresponding to the historical image, and retains the latent feature map of the historical image corresponding to the most recent several historical pages according to a preset window length.

[0063] In order to maintain the consistency of character identities and the continuity of the plot between pages during the continuous generation of picture book images, this embodiment initializes a dual-source dynamic feature memory. , It includes a core anchor point feature storage area and a historical evolution feature storage area. The core anchor point feature storage area stores latent feature maps corresponding to reference images, representing at least one of the following: facial features of the character, clothing texture features, and painting style features. The historical evolution feature storage area stores latent feature maps corresponding to historical images, representing at least one of the following: character pose change information, local appearance change information, and scene change information. The historical evolution feature storage area uses different data sources during the training and inference phases. During the training phase, the historical evolution feature storage area is constructed by importing the historical image latent feature maps extracted from the historical images of the training sample group in step (2-1). During the inference phase, the historical evolution feature storage area is dynamically updated using the historical evolution latent feature maps corresponding to the generated pages. The historical evolution feature storage area adopts a first-in, first-out (FIFO) update mechanism. When the number of historical evolution latent feature maps stored in the historical evolution feature storage area exceeds the preset window length... (This embodiment is preferred) When it is 3), delete the earliest written historical evolution potential feature map and retain the most recent one. The historical evolution potential feature map corresponding to each historical page.

[0064] (2-4) Construct an extended feature pool based on frequency domain decoupling:

[0065] Frequency domain decoupling is performed on the latent feature maps of reference images and historical images in the dual-source dynamic feature memory to extract high-frequency features used to characterize local edges, texture details, and local structural information. The high-frequency features are then processed by block segmentation and linear projection to obtain high-frequency reference visual feature sequences and high-frequency historical evolution feature sequences. These sequences are then input into the key projection layer and value projection layer of the pre-trained text conditional image generation base model to obtain reference image key-value features and historical image key-value features, respectively.

[0066] The latent feature map of the target image is processed by block segmentation and linear projection to obtain the target visual feature sequence. The target visual feature sequence is then input into the key projection layer and value projection layer of the pre-trained text conditional image generation base model to obtain the key-value features of the target image itself.

[0067] The key-value features of the target image itself, the key-value features of the reference image, and the key-value features of historical images are concatenated in the sequence dimension to construct an extended feature pool.

[0068] During the training phase, for any training sample group, the dual-source dynamic feature memory is used. The system reads feature content from the core anchor point feature storage area and the historical evolution feature storage area. When there are no historical images in the training sample group, the historical evolution feature storage area is empty, and only the reference latent feature map corresponding to the reference image is read; when there are historical images in the training sample group, the feature map closest to the target image is read from the historical evolution feature storage area. One historical latent feature map; when the number of historical images is less than At that time, all historical latent feature maps are read. In order to reduce the interference of overall pose changes, background layout changes, and large-scale composition changes in the reference image and historical images on the target image generation process, this embodiment performs frequency domain decoupling processing on the read reference latent feature maps and historical latent feature maps.

[0069] Specifically, a two-dimensional discrete wavelet transform is performed on the reference latent feature map and each historical latent feature map to transform them from the spatial domain to the frequency domain, obtaining low-frequency and high-frequency components. The low-frequency components mainly characterize the overall contour, pose layout, and large-area background structure of the image; the high-frequency components mainly characterize local edges, texture details, local structures, and painterly textures. A preset high-pass filter is used to retain the high-frequency components in the frequency domain features and suppress the low-frequency components to reduce the influence of large-scale structural information unrelated to the target page on the feature fusion process. An inverse wavelet transform is performed on the retained high-frequency components to obtain a high-frequency reference latent feature map and a high-frequency historical latent feature map. The high-frequency reference latent feature map and the high-frequency historical latent feature map are then processed by block segmentation and linear projection to obtain a high-frequency reference visual feature sequence and a high-frequency historical evolution feature sequence. The high-frequency reference visual feature sequence is input into a key projection layer and a value projection layer to obtain high-frequency reference visual key-value pair features. The high-frequency historical evolution feature sequences are input into the key projection layer and value projection layer, respectively, to obtain the high-frequency historical evolution key-value pair features. and The latent feature map of the target image is segmented and linearly projected to obtain a target visual feature sequence. This sequence is then input into a key projection layer and a value projection layer to obtain the key vector corresponding to the target image itself. Sum value vector The key features of the target image itself are then used as key features. The key features of the target image itself, the high-frequency reference visual key features, and the high-frequency historical evolution key features are sequentially concatenated to construct an extended feature pool.

[0070] When the historical evolution feature storage area is empty, the blended key vector and blended value vector in the expanded feature pool are represented as follows:

[0071] , ;in, This indicates concatenation along the feature sequence dimension.

[0072] When the historical evolution feature storage area is not empty, the blended key vector and blended value vector in the expanded feature pool are respectively:

[0073] , .

[0074] In this way, the extended feature pool simultaneously includes the features of the target image itself, the local appearance details in the reference image, and the local evolution features in the historical image, thus providing a feature foundation for subsequent cross-image feature fusion.

[0075] (2-5) Memory enhancement and attention fusion based on temporal decay bias:

[0076] The target visual feature sequence is input into the query projection layer of the pre-trained text conditional image generation base model to obtain the target image query features. A temporal decay bias is constructed based on the page distance between the historical image and the target image. The historical image features that are farther away from the target image have a lower impact on attention calculation. Memory-enhanced attention fusion is performed based on the target image query features, the key-value features in the extended feature pool, and the temporal decay bias to obtain the memory-enhanced target image generation features.

[0077] To reduce the impact of earlier historical images on the target image generation process when fusing historical evolution features, this embodiment constructs a temporal attenuation bias matrix based on the page distance between the historical images and the target image. .

[0078] For any historical image Its relationship with the target image The page distance between them is defined as: .in, This indicates the page number of the target image within the corresponding picture book sequence. This represents the page number of a historical image within the corresponding picture book sequence, and satisfies the following conditions: Set the attention bias of the historical image to the historical evolution features. .in, The preset attenuation coefficient, and Page distance The larger the value, the smaller the attention bias corresponding to the historical images, thus reducing the influence of earlier historical pages on the current target image generation process. For reference visual features and the target image's own features, their corresponding attention biases are set to 0. Therefore, a temporal decay bias matrix matching the length of the key-value sequence in the extended feature pool is constructed. The target visual feature sequence corresponding to the target image is input into the query projection layer to obtain the target query vector. Based on the target query vector Mixed bond vector Mixed value vector and timing decay bias matrix Perform memory enhancement and attention calculations: .in, This indicates the feature dimension of the attention head, and the superscript T indicates transpose. This indicates a memory-enhanced attention output that integrates high-frequency features from the reference image, high-frequency features from historical images, and features of the target image itself.

[0079] The output of the memory-enhanced attention is input to the output projection layer and residually connected with the intermediate features from the target image generation process to obtain the memory-enhanced target image generation features. These memory-enhanced target image generation features are used in the subsequent velocity field prediction process.

[0080] Through the aforementioned memory-enhanced attention fusion method, the target image can obtain stable character appearance details and painting style information from reference images during the generation process, and obtain local evolution information related to previous pages from historical images. At the same time, the interference of distant historical pages on the current page is reduced through temporal decay bias.

[0081] (2-6) Parameter optimization based on target image reconstruction loss:

[0082] The target image is trained by denoising and reconstructing the generated features of the memory-enhanced target image. The reconstruction loss is used to update the parameters of the consistency-enhanced low-rank adapter, resulting in a text conditional image generation model with cross-image feature fusion capability.

[0083] During the training phase, based on the memory-enhanced target image generation features obtained in steps (2-5), the target image is denoised and reconstructed for training to update the consistency-enhanced low-rank adapter. The parameters are specified. In this embodiment, the pre-trained text conditional image generation base model is preferably a flow-matching-based text conditional image generation model. For the target image in the training sample group, its target latent variable is obtained using a pre-trained variational autoencoder. And sample noise At time step Next, construct intermediate latent variables. The corresponding target velocity field , The time step is obtained by sampling from a preset time distribution.

[0084] The target text description, reference image features, historical image features, and intermediate latent variables are used. A consistency-enhancing low-rank adapter was inserted into the input. The text-conditional image generation base model, through the extended feature pool construction and memory-enhanced attention fusion process described in steps (2-4) and (2-5), obtains the velocity field predicted by the model. ,in, Indicates the text conditions corresponding to the target image, subscript Indicates a consistency-enhanced low-rank adapter Trainable parameters.

[0085] Based on the difference between the predicted velocity field and the target velocity field, a target image reconstruction loss is constructed: .

[0086] During training, only the consistency-enhancing low-rank adapter is updated. The parameters of the pre-trained text conditional image generation pedestal model are kept frozen. This is achieved by minimizing the target image reconstruction loss. Enhancing consistency in low-rank adapters Learning to select and blend features from reference and historical images helps maintain character consistency and page coherence.

[0087] After training, the pre-trained text conditional image generation base model, which superimposed a consistency-enhanced low-rank adapter, a dual-source dynamic feature memory, an extended feature pool, and a memory-enhanced attention fusion module, becomes a text conditional image generation model with cross-image feature fusion capabilities. It is used for candidate image generation in subsequent steps (3) and automated preference alignment driven by multimodal visual feedback.

[0088] Step (3) Construct a continuous picture book image generation model based on preference alignment.

[0089] This embodiment, based on the cross-image feature collaborative network constructed and trained in step (2), further introduces a multimodal visual feedback-driven automated preference alignment mechanism. The automated preference alignment mechanism optimizes the generation model based on the visual structural rationality, character consistency, and page coherence of the candidate generated images. This allows the model to maintain character appearance consistency and page continuity while reducing the probability of limb redundancy, local structural anomalies, facial structural anomalies, or perspective anomalies in the generated images. Step (3), while maintaining the cross-image feature collaborative capability obtained in step (2), introduces a preference alignment low-rank adapter. This further improves the visual structure quality of the generated images.

[0090] (3-1) Generate a candidate image set:

[0091] For the current picture book page, first obtain the text prompts for the current page. and reference image information Read the current state of the dual-source dynamic feature memory. Construct the composite generation conditions for the current page , Multiple candidate image samples are generated by changing random noise, sampling parameters, or repeating the sampling process, resulting in a candidate image set. During the candidate image generation process, for the same composite generation conditions... Multiple sets of candidate image samples are generated by using different random noise, different sampling parameters, or multiple sampling processes to obtain a candidate image set: ;in, This indicates that for the same composite generation conditions Number of candidate images generated Indicates the first There are 10 candidate image samples. The candidate image samples are generated under the same text description, the same reference image information, and the same dual-source dynamic feature memory. Therefore, the differences between different candidate images mainly come from random noise, sampling path, or sampling parameter differences.

[0092] Current page text prompt Used to describe the character actions, scene content, and visual style of this page; the reference image information Used to provide identity information, appearance details, and painting style information of the target character; the state of the dual-source dynamic feature memory. This includes a core anchor feature storage area and a historical evolution feature storage area. The core anchor feature storage area stores the core appearance features corresponding to the reference image, while the historical evolution feature storage area stores the evolution features corresponding to pages previously generated or historical pages. During candidate image generation, composite generation conditions are maintained. constant.

[0093] (3-2) Construct an automated preference sample dataset based on a multimodal visual evaluation model:

[0094] The candidate image set is input into a multimodal visual evaluation model (e.g., Qwen2.5-VL). Combined with preset visual evaluation instructions, the model evaluates the visual structure rationality, role consistency, and page coherence of the samples in the candidate image set, and obtains the corresponding preference evaluation scores. Based on the preference evaluation scores, positive and negative samples are selected from the candidate image samples under the same composite generation conditions. The composite generation conditions, positive samples, and negative samples are combined to construct preference triples, forming an automated preference sample dataset.

[0095] To automate the quality assessment of candidate image samples, this embodiment introduces a multimodal visual evaluation model. The multimodal visual evaluation model receives candidate image samples, reference images, historical page images or historical memory information, and preset visual evaluation instructions, and outputs a preference evaluation score corresponding to the candidate image sample. The preset visual evaluation instructions guide the multimodal visual evaluation model to evaluate candidate image samples from three aspects: visual structural rationality, role consistency, and page coherence.

[0096] Specifically, for any candidate image sample The multimodal visual evaluation model outputs an evaluation score. , , .in, The visual structure defect score is indicated. Visual structure defects include, but are not limited to, extra fingers, redundant limbs, missing limbs, abnormal joint connections, abnormal facial structures, distortion of local structures of characters, object overlap, and unreasonable perspective relationships. The larger the value, the higher the probability that there are visual structural anomalies in the candidate image. The character consistency deviation score is used to characterize the degree to which the appearance, facial features, clothing texture, or painting style of a character in a candidate image deviates from that of a reference image. The larger the value, the worse the role consistency. The page coherence deviation score represents the degree of deviation between candidate images and historical page images in terms of character state, local appearance, scene continuity, or style consistency. The larger the value, the worse the coherence between the candidate image and the preceding page.

[0097] In a preferred embodiment, a comprehensive preference evaluation score is constructed based on the above scores: ;in, The preset weighting coefficients satisfy the following conditions: This embodiment is preferred. All scores are 1. Overall Preference Rating Score The smaller the value, the more reasonable the visual structure of the candidate image, the better the consistency of its roles, and the stronger the page continuity; the overall preference evaluation score. The higher the score, the more likely the candidate image contains visual structural anomalies, character deviations, or insufficient page coherence. The overall preference score is used for evaluation. Candidate image samples are automatically filtered. Specifically, a first threshold is set. Second threshold And satisfy When candidate image samples satisfy When this happens, the candidate image sample is taken as a positive sample. When candidate image samples satisfy When this happens, the candidate image sample is used as a negative sample. In this embodiment, the first threshold The preferred value is 0.2, the second threshold. The preferred value is 0.7. Positive samples. Candidate images representing images with relatively reasonable visual structure, good character consistency, and strong page coherence; negative samples Candidate images that indicate a high degree of visual structural anomaly, poor role consistency, or weak page coherence.

[0098] Under the same compound generation conditions Positive samples and negative samples Combining to construct preference triples Multiple preference triples together constitute the automated preference sample dataset. When considering the same compound generation conditions If there are no positive or negative samples in the generated candidate images that meet the threshold or sorting conditions, discard the candidate image set, or resample the candidate images until a preference triplet that meets the conditions is obtained.

[0099] (3-3) Constructing a reference model and strategy model for parameter decoupling:

[0100] The text-conditional image generation model is used as the reference model, and the parameters of the pre-trained text-conditional image generation base model and the consistency enhancement low-rank adapter are frozen. A trainable preference-aligned low-rank adapter is superimposed on the reference model to construct a policy model. During the preference optimization stage, only the parameters of the preference-aligned low-rank adapter are updated.

[0101] The preference-aligned low-rank adapter is a parameter-efficient fine-tuning module that includes at least one set of low-rank parameter matrices and corresponding scaling coefficients. This adapter is inserted in parallel into the bypass of the original linear transformation layer of the pre-trained text conditional image generation pedestal model. While keeping the original weight matrix frozen, it generates trainable weight increments through the low-rank parameter matrices. It is trained during the preference optimization stage to adjust the generation tendency of the policy model according to automated preference samples, so as to improve the rationality of visual structure and reduce the probability of local structural anomalies.

[0102] To maintain the learned role consistency and cross-image feature fusion capabilities during preference optimization, this embodiment employs a parameter-decoupled preference alignment model structure. A text-conditional image generation model is used as the reference model. Reference Model Including a pre-trained text-conditional image generation base model and a consistency-enhancing low-rank adapter. It includes a dual-source dynamic feature memory, an expanded feature pool, and a memory-enhanced attention fusion module.

[0103] In the reference model In the middle, the backbone parameters of the pre-trained text conditional image generation base model and the consistency enhancement low-rank adapter are discussed. All parameters remain frozen; the dual-source dynamic feature memory, frequency domain decoupling module, extended feature pool, and memory-enhanced attention fusion module participate in the forward computation in accordance with step (2) to provide a reference model under composite generation conditions. The baseline generation distribution or baseline prediction results are given under the reference model. Building upon this foundation, a trainable preference-aligned low-rank adapter is further injected into the attention projection layer and feedforward network module of the generative network. Thus, a strategy model is constructed. Strategy Model Including a pre-trained text-conditional image generation base model and a consistency-enhancing low-rank adapter. Preference Alignment Low-Rank Adapter It includes a dual-source dynamic feature memory, a frequency domain decoupling module, an extended feature pool, and a memory-enhanced attention fusion module.

[0104] During the preference optimization phase, only the preference-aligned low-rank adapter is updated. The parameters, backbone parameters of the pre-trained text conditional image generation pedestal model, and consistency enhancement low-rank adapter All parameters remain frozen. Preferably, the low-rank adapter is aligned. The rank is set to 16, and the scaling factor is set to 16.

[0105] The above structure allows for parameter separation between the role consistency enhancement module and the visual preference alignment module. Specifically, the consistency enhancement low-rank adapter... To maintain consistency across image roles and page coherence, a preference-aligned low-rank adapter is used. This is used to optimize the rationality of visual structure based on automated preference samples, thereby reducing the interference of the preference optimization process on the learned role consistency features.

[0106] (3-4) The policy model is directly optimized based on the automated preference sample dataset, so that the policy model improves the preference score of positive samples and reduces the preference score of negative samples compared with the reference model. After training, a continuous picture book image generation model containing consistency enhancement low-rank adapter and preference alignment low-rank adapter is obtained.

[0107] The automated preference sample dataset constructed based on step (3-2) For strategy models Direct preference optimization is performed. In this embodiment, the pre-trained text-conditional image generation base model is a flow-matching-based generation model. Since flow-matching generation models typically do not directly output explicit normalized probability densities, this embodiment uses the velocity field prediction error difference to construct the relative preference score of the policy model relative to the reference model, which is used to approximately characterize the policy model under composite generation conditions. The degree of support for the sample images. Specifically, for any sample image... Constructing corresponding latent variables using pre-trained variational autoencoders And sample noise At time step Next, construct intermediate latent variables: ;in , indicating from the interval Random sampling from a uniform distribution yields the following target velocity field: .

[0108] Reference Model The corresponding velocity field prediction results are expressed as follows: Strategy Model The corresponding velocity field prediction results are expressed as follows: Then the sample image Under composite generation conditions The relative preference score is defined as follows: ;in, Let L be the squared norm. The relative preference score above represents: when the policy model... Relative to the reference model When the velocity field corresponding to a sample image can be predicted more accurately, the sample image has a higher relative preference score under the policy model. In one implementation, the relative preference score can be regarded as an approximate representation of the log-likelihood ratio between the policy model and the reference model.

[0109] Based on positive samples and negative samples Based on the relative preference scores, construct the direct preference optimization loss function: ;in, This represents the Sigmoid function. This represents the preference intensity adjustment coefficient. This represents the automated preference sample dataset.

[0110] The loss function is optimized by minimizing the direct preference. Make the strategy model Relative to the reference model Improve positive samples The relative preference score, and reduce the bias towards negative samples. The relative preference score.

[0111] During the optimization process, the backbone parameters of the pre-trained text conditional image generation pedestal model and the consistency enhancement low-rank adapter are... The parameters remain frozen, with only a preference for aligning low-rank adapters. The parameters participate in the backpropagation update. Through this parameter decoupling optimization method, the model can further improve the visual structural rationality of the generated image while maintaining the role consistency and cross-page coherence capabilities learned in step (2). After training, a low-rank adapter with consistency enhancement is obtained. And preference-aligned low-rank adapter A continuous picture book image generation model.

[0112] Step (4) Generate a series of children's picture book images.

[0113] (4-1) Obtain the page text sequence of a children's picture book ,in, This indicates the number of pages to be generated in the picture book. Indicates the first The text prompts for each picture book page describe the character actions, scene content, and art style on the current page. Reference image information is also retrieved. Reference image information This includes reference images used to characterize the target character's identity, appearance details, clothing textures, and painting style, or reference latent feature maps encoded from reference images.

[0114] (4-2) Based on reference image information The dual-source dynamic feature memory is initialized by writing the reference latent feature map corresponding to the reference image into the core anchor feature storage area. Before generating the first picture book page, the historical evolution feature storage area is empty. When generating subsequent picture book pages, the historical evolution feature storage area is used to store the historical latent feature maps corresponding to the generated pages.

[0115] (4-3) Generate the current picture book images in page order:

[0116] For the currently to be generated, the first Each picture book page reads the text prompts for the current page. Reference image information and the current state of the dual-source dynamic feature memory. Construct composite generation conditions: Composite generation conditions Input the target continuous picture book image generation model. Following the frequency domain decoupling, extended feature pool construction, and memory-enhanced attention fusion process described in step (2), the model extracts character appearance and painting style information from the reference image, and extracts state evolution information related to previous pages from the historical evolution feature storage area to generate the current page image: , This represents the continuous picture book image generation model trained in step (3). Indicates the generated first Images from a picture book page.

[0117] (4-4) Encode the current page image to obtain the current page potential feature map and write it into the historical evolution feature storage area; when the number of historical page features stored in the historical evolution feature storage area exceeds the preset window length, delete the earliest written historical page potential feature map and retain the historical potential feature maps corresponding to the most recently generated pages; repeat the current page image generation and historical evolution feature storage area update process according to the page order until a complete continuous picture book image sequence is obtained.

[0118] (4-4) Update the historical evolution feature storage area and continue generating:

[0119] Generating the current page image Then, the current page image is encoded to obtain the current page potential feature map, and it is written into the historical evolution feature storage area for subsequent page generation.

[0120] When the number of historical page features stored in the historical evolution feature storage area exceeds the preset window length At that time, the earliest written historical page potential feature map is deleted according to the first-in-first-out rule, and only the most recent one is retained. The latent feature map corresponding to each generated page.

[0121] Repeat steps (4-3) and (4-4) in page order until a complete sequence of consecutive picture book images is generated:

[0122] .

[0123] In this way, the model can maintain the consistency of the reference characters and the painting style during continuous generation, and enhance the coherence between pages by utilizing information from recent historical pages.

[0124] As a preferred engineering implementation of the above method, the model training and inference process in this embodiment is deployed on an industrial-grade high-performance computing server. The hardware configuration preferably includes an NVIDIA H800 GPU accelerator card cluster or equivalent computing power. The operating system is Ubuntu 20.04, the deep learning framework is PyTorch 2.0, and the model architecture is further customized based on the Diffusers open-source library. For network training hyperparameter settings, the AdamW optimizer is uniformly used, and the initial learning rate is set to [value missing]. The mixed precision training mechanism is enabled throughout the process to greatly accelerate the convergence throughput of large models and optimize memory usage while ensuring the numerical stability of gradient calculation.

[0125] Embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, portable hard drives, USB flash drives, cloud computing platform storage, and solid-state drives, etc.) containing computer-usable program code.

[0126] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0127] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0128] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0129] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features, and such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for generating continuous children's picture book images based on dynamic memory enhancement, characterized in that, Includes the following steps: Step (1) Collect raw image data, and obtain image-text pair dataset after preprocessing. The raw image data is children's picture book image data. Step (2) Construct a text conditional image generation model with cross-image feature fusion capability; (2-1) Insert a consistency enhancement low-rank adapter into the pre-trained text conditional image generation base model, keep the backbone parameters of the pre-trained text conditional image generation base model frozen, and update only the parameters of the consistency enhancement low-rank adapter during the training phase. (2-2) Constructing latent feature maps for multi-source images: The image-text pair dataset is organized into training sample groups according to role association and page order. Each training sample group includes a target image-text pair, one or more reference image-text pairs, and multiple historical image-text pairs. The target image, reference image, and historical image in the target image-text pair, reference image-text pair, and historical image-text pair are encoded using a pre-trained variational autoencoder to map the target image, reference image, and historical image to the latent space to obtain the corresponding latent feature maps. (2-3) Construct a dual-source dynamic feature memory: The dual-source dynamic feature memory includes a core anchor feature storage area and a historical evolution feature storage area; wherein, the core anchor feature storage area is used to store the latent feature map corresponding to the reference image, and the historical evolution feature storage area is used to store the latent feature map corresponding to the historical image, and the latent feature map of the historical image corresponding to the most recent historical page is retained according to a preset window length. (2-4) Construct an extended feature pool based on frequency domain decoupling: Frequency domain decoupling is performed on the latent feature maps of reference images and historical images in the dual-source dynamic feature memory to extract high-frequency features that characterize local edges, texture details, and local structural information. The high-frequency features are then segmented and linearly projected to obtain high-frequency reference visual feature sequences and high-frequency historical evolution feature sequences. These sequences are then input into the key projection layer and value projection layer of the pre-trained text conditional image generation base model to obtain reference image key-value features and historical image key-value features, respectively. The latent feature map of the target image is processed by block segmentation and linear projection to obtain the target visual feature sequence. The target visual feature sequence is then input into the key projection layer and value projection layer of the pre-trained text conditional image generation base model to obtain the key-value features of the target image itself. The key features of the target image itself, the key features of the reference image, and the key features of the historical images are concatenated in the sequence dimension to construct an extended feature pool; (2-5) Memory enhancement and attention fusion based on temporal decay bias: The target visual feature sequence is input into the query projection layer of the pre-trained text conditional image generation base model to obtain the target image query features; a temporal decay bias is constructed based on the page distance between the historical image and the target image, and the historical image features that are farther away from the target image have a lower impact on the attention calculation; memory-enhanced attention fusion is performed based on the target image query features, the key-value features in the extended feature pool, and the temporal decay bias to obtain the memory-enhanced target image generation features. (2-6) Parameter optimization based on target image reconstruction loss: The target image is trained by denoising and reconstructing the generated features of the memory-enhanced target image. The reconstruction loss is used to update the parameters of the consistency-enhanced low-rank adapter. After training, the pre-trained text conditional image generation base model, which superimposes the consistency-enhanced low-rank adapter, the dual-source dynamic feature memory, the extended feature pool, and the memory-enhanced attention fusion module, is a text conditional image generation model with cross-image feature fusion capability. Step (3) Construct a continuous picture book image generation model based on preference alignment; (3-1) Generate a candidate image set: For the current picture book page, first obtain the text prompts and reference image information of the current page, read the state of the dual-source dynamic feature memory at the current moment, and construct the composite generation conditions of the current page; by changing the random noise, sampling parameters, or repeatedly executing the sampling process, generate multiple candidate image samples to obtain a candidate image set; (3-2) Construct an automated preference sample dataset based on a multimodal visual evaluation model: The candidate image set is input into the multimodal visual evaluation model. Combined with the preset visual evaluation instructions, the model evaluates the visual structure rationality, role consistency and page coherence of the samples in the candidate image set and obtains the corresponding preference evaluation score. Based on the preference evaluation score, positive and negative samples are selected from the candidate image samples under the same composite generation conditions. The composite generation conditions, positive samples and negative samples are combined to construct preference triplets to form an automated preference sample dataset. (3-3) The text conditional image generation model is used as the reference model, and the parameters of the pre-trained text conditional image generation base model and the consistency enhancement low-rank adapter are frozen; a trainable preference alignment low-rank adapter is superimposed on the reference model as a policy model, and in the preference optimization stage, only the parameters of the preference alignment low-rank adapter are updated. (3-4) Based on the automated preference sample dataset, the policy model is directly optimized to improve the preference score of the policy model for positive samples and reduce the preference score of the policy model for negative samples compared with the reference model. After training, a continuous picture book image generation model including consistency enhancement low-rank adapter and preference alignment low-rank adapter is obtained. Step (4) Generate continuous children's picture book images; (4-1) Obtain the page text sequence and reference image information of children's picture books; the page text sequence includes multiple page text prompts, each of which is used to describe the character's actions, scene content and picture style in the corresponding picture book page; the reference image information is used to provide the target character's identity information, appearance details and drawing style information; (4-2) Initialize the dual-source dynamic feature memory based on the reference image information and write the reference latent feature map corresponding to the reference image into the core anchor point feature storage area; (4-3) For the current picture book page to be generated, read the text prompts, reference image information and the current state of the dual-source dynamic feature memory bank, and construct composite generation conditions; input the composite generation conditions into the continuous picture book image generation model to generate the current page image; (4-4) Encode the current page image to obtain the current page potential feature map and write it into the historical evolution feature storage area; repeat the current page image generation and historical evolution feature storage area update process according to the page order until a complete continuous picture book image sequence is obtained.

2. The method for generating continuous children's picture book images based on dynamic memory enhancement as described in claim 1, characterized in that: In step (1), the preprocessing is specifically as follows: Calculate the narrative confidence score for each original image data, remove non-narrative pages with confidence scores below a preset threshold, and obtain usable image data; The text region in the available image data is located using an optical character detection model, a binary mask is generated, the available image data and the binary mask are input into the image redrawing model, the background texture occluded by the text is reconstructed using global context information, and a clean image is output. The pixel resolution and aspect ratio of a clean image are detected, and distorted images with resolutions below the preset minimum size or aspect ratios outside the preset range are removed to obtain high-quality images. The Hamming distance of structured features between any two high-quality images is calculated using the perceptual hashing algorithm. Approximately duplicate images with similarity values ​​less than a similarity threshold are then removed to obtain the final image data. By using a multimodal large language model to perform visual understanding on the final image data, a refined text description containing character appearance, actions, and scene details is generated, resulting in an image-text pair dataset.

3. The method for generating continuous children's picture book images based on dynamic memory enhancement as described in claim 1, characterized in that: In step (2), the consistency-enhanced low-rank adapter is a parameter-efficient fine-tuning module, which includes at least one set of low-rank parameter matrices and corresponding scaling coefficients; The adapter is inserted in parallel into the bypass of the original linear transformation layer of the pre-trained text conditional image generation pedestal model. While keeping the original weight matrix frozen, it generates trainable weight increments through the low-rank parameter matrix to supplement and correct the feature transformation results of the pre-trained text conditional image generation pedestal model.

4. The method for generating continuous children's picture book images based on dynamic memory enhancement as described in claim 1, characterized in that: In step (2), the target image-text pair includes the target image and its corresponding text description, which is used to provide the target sample information to be generated; the reference image-text pair includes the reference image and its corresponding text description, which is used to provide the character identity information and painting style information; the historical image-text pair includes the historical image and its corresponding text description, which is used to provide the state evolution information of the target page before.

5. The method for generating continuous children's picture book images based on dynamic memory enhancement as described in claim 1, characterized in that: In step (2), the historical evolution feature storage area adopts a first-in-first-out update mechanism, retaining only the historical evolution potential feature map corresponding to the most recent historical page of a set length.

6. The method for generating continuous children's picture book images based on dynamic memory enhancement as described in claim 1, characterized in that: In step (3), the current page text prompt is used to describe the character's actions, scene content and screen style on the page; the reference image information is used to provide the target character's identity information, appearance details and painting style information; during the candidate image generation process, the composite generation conditions remain unchanged and the dual-source dynamic feature memory state remains unchanged.

7. The method for generating continuous children's picture book images based on dynamic memory enhancement as described in claim 1, characterized in that: In step (3), the preference-aligned low-rank adapter is a parameter-efficient fine-tuning module, which includes at least one set of low-rank parameter matrices and corresponding scaling coefficients. The adapter is inserted in parallel into the bypass of the original linear transformation layer of the pre-trained text conditional image generation pedestal model. While keeping the original weight matrix frozen, it generates trainable weight increments through the low-rank parameter matrix. It is trained in the preference optimization stage to adjust the generation tendency of the strategy model according to the automated preference samples, so as to improve the rationality of visual structure and reduce the probability of local structural anomalies.