A novel method and system for generating a quadrilateral continuous graph based on image inpainting
Patent Information
- Application Number
- CN202611035969.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-13
- Publication Date
- 2026-10-09
AI Technical Summary
[0003]然而,对于纹理结构、颜色分布及图案内容较复杂的RGB图案图像,现有处理方式难以充分利用未修复区域所包含的整体视觉语义信息,导致接缝修复内容与原图案之间存在视觉特征不协调的问题
本发明通过对待处理的RGB图案图像进行水平循环滚动和垂直循环滚动,将原始图像的左右边界接缝及上下边界接缝依次移动至图像内部,并利用竖直条形二值掩膜对水平接缝待修复区域进行明确限定,使边界接缝能够转化为内部局部区域的图像修复问题;通过掩码自编码器提取水平带掩膜图像中未屏蔽区域的视觉特征,获得能够反映原图案内容和纹理分布的水平视觉语义特征,再利用CLIP对齐子模块和T5对齐子模块进行双路特征转换,使所述视觉语义信息能够以不同维度的条件信息共同参与整流流扩散模型的去噪过程,降低修复内容与原图案在语义、纹理及局部结构上的不协调现象。
Smart Images

Figure CN122888043A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a novel method and system for generating four-sided continuous images based on image inpainting. Background Technology
[0002] A four-way continuous image refers to an image format where the image content remains continuous when patterns are repeatedly stitched together in the horizontal and vertical directions. During pattern design and image processing, the left and right boundaries and top and bottom boundaries of the original pattern need to be processed to reduce noticeable seams when repeated. Current four-way continuous image generation processes typically handle boundary areas through image displacement, edge stitching, or local image inpainting. Image displacement can move boundary seams into the image, but further content filling is still required in the seam areas; local image inpainting mainly generates filling content based on the pixel or texture information surrounding the area to be repaired.
[0003] However, for RGB pattern images with complex textures, color distributions, and pattern content, existing processing methods struggle to fully utilize the overall visual semantic information contained in the unrepaired areas, leading to visual inconsistencies between the repaired seam content and the original pattern. Furthermore, after latent space encoding, denoising, and decoding, boundary transition marks easily form between the repaired and original image areas, affecting the continuity when cyclically splicing left and right or top and bottom boundaries. In addition, existing processing workflows often require separate image repair processes when generating multiple repair results, making it difficult to simultaneously perform latent space denoising on multiple pattern variants using batch data. Therefore, how to combine pattern visual semantics to repair the seam areas after cyclic scrolling, reduce boundary differences in the repaired areas, and simultaneously achieve batch generation of multiple four-sided continuous image variants has become a problem that needs to be solved in existing four-sided continuous image generation technologies.
[0004] Therefore, how to provide a novel method and system for generating quadrangular continuous graphs based on image inpainting is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] One objective of this invention is to propose a novel method and system for generating quadrangular continuous graphs based on image inpainting. This invention integrates visual semantic extraction, dual-path conditional alignment and rectified flow latent space repair, and combines it with Laplacian pyramid fusion to achieve automatic generation of quadrangular continuous graphs, thereby improving the consistency of seam repair, the naturalness of boundary transitions, and the efficiency of batch generation of multiple variants.
[0006] A novel method for generating a tetrahedral continuous graph based on image inpainting according to an embodiment of the present invention includes the following steps: Step 1: Obtain the RGB pattern image to be processed, assign different random seeds according to the preset variant generation quantity, perform horizontal cyclic scrolling on the RGB pattern image, and generate a vertical strip binary mask to obtain a horizontal masked image; Step 2: Extract visual features from the horizontal masked image using a mask autoencoder to generate horizontal visual semantic features; Step 3: Input the horizontal visual semantic features into the dual-path alignment submodule for feature transformation to obtain horizontal visual condition information including the horizontal CLIP condition vector and the horizontal T5 condition sequence; Step 4: Input the horizontal masked image into the variational autoencoder to compress the horizontal masked image into the latent space, perform downsampling and rearrangement, and initialize Gaussian noise with a random seed to generate horizontal batch noise input; Step 5: Input the horizontal batch noise input and horizontal visual condition information into the rectified flow diffusion model for iterative denoising to generate N horizontal repair latent codes; Step 6: Input the N horizontal insulated latent codes into the standard variational autoencoder decoder for batch decoding, and generate N horizontal boundary fusion images through Laplacian pyramid fusion and boundary smoothing processing; Step 7: Perform reverse loop rolling on the N horizontal boundary fusion images respectively to generate N horizontal seamless images; Step 8: Perform vertical cyclic scrolling on the N horizontal seamless images respectively, repeating steps one to seven to obtain a four-sided continuous image.
[0007] Optionally, step one specifically includes: Obtain the RGB pattern image to be processed, determine the image width and image height of the RGB pattern image, and determine the number N of tetrahedral continuous image variants to be generated according to the preset number of variants generated. Assign different random seeds to each tetrahedral continuous image variant to generate the pattern image processing task. For each pattern image processing task, the RGB pattern image is cyclically scrolled horizontally, and each pixel in the image is cyclically shifted horizontally by 1 / 2 of the image width to generate a horizontally scrolling image. A vertical strip binary mask is generated based on the horizontal center position of the horizontally looping image. The area covered by the vertical strip binary mask is determined as the horizontal seam to be repaired, and a horizontal masked image is generated.
[0008] Optionally, step two specifically includes: The horizontal masked image is scaled to 256×256 pixels and normalized according to the mean and standard deviation of the ImageNet dataset. Clear the pixel values of the region corresponding to the vertical bar binary mask in the normalized image to zero, and generate the horizontally visible region image; The horizontally visible area image is input into a mask autoencoder. Visual features are extracted through the encoder and decoder of the mask autoencoder. The intermediate features of the 6th layer decoder of the mask autoencoder are extracted to generate horizontal visual semantic features with a dimension of N×512.
[0009] Optionally, step three specifically includes: The dual-path alignment submodule includes a CLIP alignment submodule and a T5 alignment submodule; The horizontal visual semantic features are input into the CLIP alignment submodule and the T5 alignment submodule, respectively. The CLIP alignment submodule converts the horizontal visual semantic features into a horizontal CLIP conditional vector with an output dimension of 768, and the T5 alignment submodule converts the horizontal visual semantic features into a horizontal T5 conditional sequence with an output dimension of 4096. The horizontal CLIP condition vector is used as the first visual condition information, the horizontal T5 condition sequence is used as the second visual condition information, and the first and second visual condition information are combined into the horizontal visual condition information.
[0010] Optionally, step four specifically involves: The horizontal masked image is input into a variational autoencoder with float32 precision, and the horizontal masked image is compressed to a latent space with 16 channels and a spatial size downsampled by 8 times to generate a horizontal image latent code. The vertical strip binary mask is downsampled in an 8×8 block manner and rearranged in a 2×2 packing manner. The processed mask data is then spliced with the horizontal image latent code along the channel dimension to generate a horizontal joint conditional input. Each independent Gaussian noise with the same latent coding shape as the horizontal image is initialized according to the random seed corresponding to each four-dimensional continuous graph variant. The N Gaussian noises and the corresponding N horizontal joint conditional inputs are then concatenated along the batch dimension to generate the horizontal batch noise input.
[0011] Optionally, step five specifically includes: The horizontal batch noise input and horizontal visual condition information are input into the rectified flow diffusion model for iterative denoising to generate a 50-step rectified flow time step schedule. Time-shift scheduling is used to adjust the denoising rhythm for images of different resolutions; A 50-step iterative denoising process is performed with a guiding scale of 30. Forward calculation is performed through a diffusion transformer at each denoising time step. The denoising guiding direction is determined by the horizontal CLIP condition vector and the horizontal T5 condition sequence. Batch latent space repair is performed on the horizontal seam repair areas corresponding to N quadrature continuous graph variants. The denoised latent codes are restored from sequence form to spatial form, generating N horizontally repaired latent codes.
[0012] Optionally, step six specifically includes: The N horizontal restoration latent codes are input into a standard variational autoencoder decoder for batch decoding to generate N horizontal restoration RGB images; An L-layer Gaussian pyramid is constructed based on each horizontally repaired RGB image, the corresponding horizontally looping image, and the vertical strip binary mask. The vertical strip binary mask is simultaneously downsampled by Gaussian, and the Laplacian pyramids corresponding to the horizontally repaired RGB image and the horizontally looping image are generated based on the adjacent Gaussian pyramid layers. In each pyramid layer, the Laplacian layer of the horizontally repaired RGB image and the Laplacian layer of the horizontally looping image are weighted and fused according to the corresponding layer mask. The fused Laplacian pyramid is then reconstructed layer by layer from coarse to fine to generate N horizontal boundary fused images.
[0013] Optionally, step seven specifically includes: Perform reverse loop scrolling along the horizontal direction on each of the N horizontal boundary fused images; The reverse loop rolling distance is negative 1 / 2 of the image width, until the horizontal boundary fused image is restored to the original image coordinate position, generating N horizontal seamless images; The left and right boundaries of seamless images at each level form a continuous stitching relationship.
[0014] A novel four-sided continuous graph generation system based on image inpainting according to an embodiment of the present invention includes the following modules: The horizontal mask processing module is used to acquire the RGB pattern image to be processed, allocate different random seeds according to the preset variant generation quantity, perform horizontal cyclic scrolling on the RGB pattern image, and generate a vertical strip binary mask to obtain a horizontal masked image. The visual semantic extraction module is used to extract visual features from the horizontal masked image through a mask autoencoder to generate horizontal visual semantic features. The dual-path condition alignment module is used to input the horizontal visual semantic features into the dual-path alignment submodule for feature transformation, and obtain horizontal visual condition information including the horizontal CLIP condition vector and the horizontal T5 condition sequence. The latent space coding module is used to input the horizontal band mask image into the variational autoencoder, compress the horizontal band mask image into the latent space, perform downsampling and rearrangement, and initialize Gaussian noise with a random seed to generate horizontal batch noise input. The rectified flow denoising module is used to input the horizontal batch noise input and the horizontal visual condition information into the rectified flow diffusion model for iterative denoising, generating N horizontal repair latent codes; The decoding and fusion module is used to batch decode the N horizontal repair latent coding inputs to the standard variational autoencoder decoder, and generate N horizontal boundary fusion images through Laplacian pyramid fusion and boundary smoothing processing. The horizontal seamless generation module is used to perform reverse cyclic rolling on N horizontal boundary fusion images respectively to generate N horizontal seamless images; The vertical seamless generation module is used to perform vertical cyclic scrolling on N horizontal seamless images respectively, repeating steps one to seven to obtain a four-sided continuous image.
[0015] The beneficial effects of this invention are: This invention uses horizontal and vertical cyclic scrolling to move the left and right boundary seams and top and bottom boundary seams of the original image into the image interior. A vertical strip binary mask is used to clearly define the horizontal seam repair area, transforming the boundary seam problem into an image repair problem of an internal local area. A mask autoencoder extracts the visual features of the unmasked areas in the horizontally masked image, obtaining horizontal visual semantic features that reflect the original pattern content and texture distribution. Then, a CLIP alignment submodule and a T5 alignment submodule are used for dual-path feature transformation, allowing the visual semantic information to participate in the denoising process of the rectified flow diffusion model with conditional information of different dimensions, reducing the inconsistency between the repaired content and the original pattern in terms of semantics, texture, and local structure.
[0016] By compressing the image into the latent space using a variational autoencoder and combining it with downsampling, rearrangement, and channel stitching of mask data, the rectified flow diffusion model can perform conditionally constrained latent space repair on the horizontal seam to be repaired area. Simultaneously, batch noise input is constructed based on different random seeds, obtaining multiple horizontal repair latent codes in a single batch processing step, improving the generation efficiency of multiple quadruple continuous image variants. After decoding with a standard variational autoencoder, Gaussian and Laplacian pyramids are constructed. At different levels, the repaired image and the horizontally looping image are weighted and fused according to the mask, and reconstructed layer by layer from coarse to fine, reducing boundary differences and transition traces between the repaired area and the original image area. After restoring the original image coordinates through reverse looping, the left and right boundaries of the image are made into a continuous stitching relationship. Then, vertical corresponding processing is performed on the horizontally seamless image, ultimately obtaining a quadruple continuous image that can be continuously stitched in both the horizontal and vertical directions. This improves the consistency of seam repair, the naturalness of boundary transitions, and the batch generation capability when automatically generating quadruple continuous images from complex RGB patterns. Attached Figure Description
[0017] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is an overall flowchart of a novel four-sided continuous graph generation method based on image inpainting proposed in this invention; Figure 2 This is a schematic diagram of the structure of a novel four-sided continuous graph generation system based on image inpainting proposed in this invention. Detailed Implementation
[0018] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0019] refer to Figure 1 A novel method for generating tetrahedral continuous graphs based on image inpainting includes: Step 1: Obtain the RGB pattern image to be processed, assign different random seeds according to the preset variant generation quantity, perform horizontal cyclic scrolling on the RGB pattern image, and generate a vertical strip binary mask to obtain a horizontal masked image; Step 2: Extract visual features from the horizontal masked image using a mask autoencoder to generate horizontal visual semantic features; Step 3: Input the horizontal visual semantic features into the dual-path alignment submodule for feature transformation to obtain horizontal visual condition information including the horizontal CLIP condition vector and the horizontal T5 condition sequence; Step 4: Input the horizontal masked image into the variational autoencoder to compress the horizontal masked image into the latent space, perform downsampling and rearrangement, and initialize Gaussian noise with a random seed to generate horizontal batch noise input; Step 5: Input the horizontal batch noise input and horizontal visual condition information into the rectified flow diffusion model for iterative denoising to generate N horizontal repair latent codes; Step 6: Input the N horizontal insulated latent codes into the standard variational autoencoder decoder for batch decoding, and generate N horizontal boundary fusion images through Laplacian pyramid fusion and boundary smoothing processing; Step 7: Perform reverse loop rolling on the N horizontal boundary fusion images respectively to generate N horizontal seamless images; Step 8: Perform vertical cyclic scrolling on the N horizontal seamless images respectively, repeating steps one to seven to obtain a four-sided continuous image.
[0020] In this embodiment, step one specifically includes: Read the RGB pattern image to be processed, obtain the number of pixel columns along the horizontal direction and the number of pixel rows along the vertical direction of the RGB pattern image, and determine the number of pixel columns and the number of pixel rows as the image width and image height, respectively; Read the preset number of variants generated, determine the preset number of variants generated as the number of four-sided continuous pattern variants N, establish a one-to-one correspondence between each four-sided continuous pattern variant and a random seed, configure different random seeds for different four-sided continuous pattern variants, associate the RGB pattern image, image width, image height, number of four-sided continuous pattern variants N and each random seed, and generate a pattern image processing task. In each pattern image processing task, the pixel column of the RGB pattern image is used as the horizontal cyclic displacement object, and each pixel is moved horizontally by 1 / 2 of the image width; Pixels that have moved beyond one side of the RGB pattern image are cyclically written to the other side of the image according to the corresponding displacement distance, so that the pixels corresponding to the original left boundary and the pixels corresponding to the original right boundary are moved to the horizontal center of the image, while keeping the vertical position of each pixel unchanged, thus generating a horizontally cyclically scrolling image. The horizontal center position is determined based on the image width of the horizontally scrolling image. The horizontal center position is used as the center position of the vertical strip binary mask. The vertical strip binary mask is set through the horizontally scrolling image in the vertical direction. The width of the vertical strip binary mask is 256 pixels by default. The pixel area covered by the vertical strip binary mask is determined as the horizontal seam to be repaired, thus obtaining a horizontal masked image.
[0021] In this embodiment, step two specifically includes: The horizontal strip mask image is scaled up, and the number of pixels in the horizontal direction and the number of pixels in the vertical direction of the horizontal strip mask image are adjusted to 256 respectively to generate a 256×256 pixel horizontal standard image. The vertical strip binary mask is then synchronously mapped to the pixel position corresponding to the horizontal standard image. The color channels of the horizontal standard image are normalized according to the mean of the red channel (0.485), the mean of the green channel (0.456), and the mean of the blue channel (0.406) in the ImageNet dataset, as well as the standard deviations of the red channel (0.229), the green channel (0.224), and the blue channel (0.225), respectively, to generate a horizontally normalized image. The mask pixel positions in the horizontally normalized image are determined based on the vertical strip binary mask after synchronous mapping. The red, green, and blue channel pixel values corresponding to each mask pixel position are set to 0 respectively, and the normalized pixel values corresponding to each pixel position outside the vertical strip binary mask are retained to generate a horizontally visible area image. The horizontal visible region image is divided into multiple image blocks, and each image block is arranged according to its spatial position in the horizontal visible region image to generate a horizontal image block sequence containing N image blocks. The horizontal image block sequence is then input into a mask autoencoder trained by fine-tuning the image inpainting mask distribution. The mask autoencoder is a self-supervised pre-trained model based on a visual transformer architecture, which performs the function of visual semantic condition extraction in this system. The encoder of the mask autoencoder sets the mask value corresponding to the horizontal seam to be repaired area and the mask value corresponding to the other areas to different binary states, and clears the pixel value of the horizontal seam to be repaired area to zero according to the vertical strip binary mask, while retaining the image pixels of the area outside the vertical strip binary mask. Visual representation extraction is performed on image blocks in the horizontal image block sequence that are not covered by the vertical strip binary mask to generate horizontal visible region coding features. The horizontal visible region coding features are then input into the decoder of the mask autoencoder. The decoder processes the visual representation corresponding to each image block layer by layer. The intermediate features output by the 6th layer decoder of the mask autoencoder for N image blocks are obtained. The intermediate features corresponding to each image block are represented as 512-dimensional features. The N 512-dimensional features are combined according to the order of each image block in the horizontal image block sequence to generate horizontal visual semantic features with a dimension of N×512.
[0022] In this embodiment, step three specifically includes: Read each of the 512-dimensional features in the horizontal visual semantic features according to the arrangement order of each image block in the horizontal image block sequence, match each of the 512-dimensional features with the corresponding image block arrangement position, generate a horizontal visual semantic feature sequence, and input the horizontal visual semantic feature sequence into the CLIP alignment submodule and the T5 alignment submodule respectively. The CLIP alignment submodule performs dimensional alignment on the 512-dimensional features in the horizontal visual semantic feature sequence, transforms the visual semantic information corresponding to each image block into the CLIP conditional feature dimension, and combines the transformed features according to the arrangement order of the horizontal visual semantic feature sequence to generate a horizontal CLIP conditional vector with an output dimension of 768. The T5 alignment submodule performs dimensional alignment on the 512-dimensional features in the horizontal visual semantic feature sequence, transforms the visual semantic information corresponding to each image block into the T5 conditional feature dimension, maintains the arrangement correspondence between the visual semantic information of each image block, and organizes the transformed features according to the arrangement order of the horizontal visual semantic feature sequence to generate a horizontal T5 conditional sequence with an output dimension of 4096. The horizontal CLIP condition vector and the horizontal T5 condition sequence are associated with the same horizontal masked image. The horizontal CLIP condition vector is used as the first visual condition information, the horizontal T5 condition sequence is used as the second visual condition information, and the first visual condition information and the second visual condition information are combined into horizontal visual condition information.
[0023] In this embodiment, step four specifically includes: The data precision of the horizontal mask image is converted to float32 precision, and the converted horizontal mask image is input into the variational autoencoder encoder. The variational autoencoder encoder performs spatial compression on the horizontal mask image, reducing the horizontal and vertical spatial dimensions to 1 / 8 of the corresponding original spatial dimensions (spatial downsampling by 8 times), and sets the number of channels of the encoding result to 16 to generate the horizontal image latent code. The vertical strip binary mask is read, and the vertical strip binary mask is divided into blocks according to the 8×8 pixel region, so that each 8×8 pixel region corresponds to the latent space position in the latent coding of the horizontal image, and downsampling mask data is generated according to the mask state of each block region. The downsampled mask data is grouped according to adjacent 2×2 latent space positions, and the mask data corresponding to each group of 2×2 latent space positions is rearranged to the channel direction according to a fixed spatial arrangement order to generate packaged mask data. Based on the spatial correspondence between the horizontal image latent coding and the packed mask data, the horizontal and vertical spatial positions of the two are mapped, and the packed mask data and the horizontal image latent coding are spliced along the channel dimension so that each latent spatial position simultaneously contains the corresponding image latent coding data and mask state data, generating a horizontal joint conditional input. Read the random seed corresponding to each quadruple continuous graph variant, initialize the Gaussian noise generation process with each random seed, generate an independent Gaussian noise with the same number of channels and spatial size as the latent coding of the horizontal image for each quadruple continuous graph variant, so that the Gaussian noises corresponding to different random seeds are independent of each other, and generate N Gaussian noises. The horizontal joint condition input is input and a correspondence is established with N quadruple continuum variants respectively to generate N sets of horizontal joint condition inputs; According to the correspondence order between each quadruple continuous graph variant and the random seed, N Gaussian noises are arranged along the batch dimension, and N horizontal joint condition inputs are arranged along the batch dimension in the same correspondence order, so that the Gaussian noise at the same batch position corresponds to the same quadruple continuous graph variant, thus generating horizontal batch noise input.
[0024] In this embodiment, step five specifically includes: The horizontal batch noise input, the horizontal CLIP condition vector, and the horizontal T5 condition sequence are input into the rectified flow diffusion model; A rectified stream time step sequence containing 50 denoising time steps is generated based on the spatial dimensions corresponding to the latent coding of the horizontal image. The iteration order of the latent coding from the Gaussian noise state to the image latent coding state is determined according to the order of each denoising time step in the rectified flow time step sequence. The 50 denoising time steps are adjusted by time shift scheduling. The distribution of each denoising time step in the iteration process is changed according to the spatial resolution corresponding to the horizontal image latent coding, and a 50-step rectified flow time step schedule is generated. At the current denoising time step, the current latent code, horizontal joint condition input, horizontal CLIP condition vector, and horizontal T5 condition sequence are input into the diffusion transformer for forward computation. The horizontal CLIP condition vector provides 768-dimensional visual conditions, and the horizontal T5 condition sequence provides 4096-dimensional visual condition sequences. The two together form the denoising guidance direction corresponding to the current denoising time step. The guiding scale is set to 30, the latent code of the current denoising time step is updated according to the denoising guiding direction, and the updated latent code is used as the input of the next denoising time step. According to the 50-step rectified flow time step scheduling, the forward calculation of the diffusion converter and the latent code update are repeated sequentially until the 50-step iterative denoising is completed. The latent codes corresponding to each batch position are used to perform latent space repair on the horizontal seam to be repaired area under the constraints of the horizontal joint condition input and the horizontal visual condition information, generating N denoised horizontal latent codes. According to the channel arrangement and spatial position arrangement of the horizontal image latent coding, the sequence data of each denoised horizontal latent coding are rearranged to restore the sequence form of the latent coding to the spatial form corresponding to the channel dimension, horizontal spatial dimension and vertical spatial dimension. The spatial form latent coding corresponding to each four-dimensional continuous image variant is extracted according to the batch position to generate N horizontal repair latent codes.
[0025] In this embodiment, step six specifically includes: The N horizontal restoration latent codes are input into the standard variational autoencoder decoder in batch order. The standard variational autoencoder decoder performs batch decoding on each horizontal restoration latent code, maps each horizontal restoration latent code from latent space features to pixel space, and outputs N horizontal restoration RGB images according to the correspondence of each four-sided continuous graph variant. Read the horizontally repaired RGB image, the horizontally scrolling image, and the vertical bar binary mask corresponding to each quadruple continuous image variant, and maintain the one-to-one correspondence between the three; for each quadruple continuous image variant, use the horizontally repaired RGB image as the first input image, the horizontally scrolling image as the second input image, and the vertical bar binary mask as the fusion region indication information to construct the corresponding boundary fusion processing data; The first input image and the second input image are respectively subjected to layer-by-layer smoothing and layer-by-layer downsampling to generate an L-layer Gaussian pyramid corresponding to the first input image and an L-layer Gaussian pyramid corresponding to the second input image. Perform synchronous Gaussian downsampling processing on the vertical strip binary mask corresponding to the Gaussian pyramid level to generate an L-layer mask pyramid, so that each layer of mask and the corresponding layer of Gaussian pyramid have the same spatial resolution. Based on the adjacent Gaussian pyramid layers corresponding to the first input image, a Laplacian pyramid corresponding to the first input image is generated; based on the adjacent Gaussian pyramid layers corresponding to the second input image, a Laplacian pyramid corresponding to the second input image is generated; wherein, each Laplacian pyramid layer represents the inter-layer difference information between the current layer image of the corresponding Gaussian pyramid and the adjacent next layer image; At each pyramid layer, the mask data of the corresponding layer is read, and the Laplacian layer of the first input image and the Laplacian layer of the second input image are weighted and fused according to the mask of the layer, so that the horizontal seam to be repaired area indicated by the mask adopts the Laplacian layer information of the first input image, and the area not indicated by the mask adopts the Laplacian layer information of the second input image, thereby generating the fused Laplacian image of each layer. The Laplacian images of each layer are reconstructed layer by layer in order from coarse to fine. The coarsest layer fusion result is used as the initial reconstruction result. The reconstruction result of the current layer is restored to the next higher resolution and then combined with the corresponding next layer fusion Laplacian image until the reconstruction of the finest layer is completed. The horizontal boundary fusion images corresponding to each tetrahedral continuous image variant are generated and N horizontal boundary fusion images are generated.
[0026] In this embodiment, step seven specifically includes: Read N horizontal boundary fusion images respectively, and obtain the image width corresponding to each horizontal boundary fusion image, maintaining the one-to-one correspondence between each horizontal boundary fusion image and the corresponding tetrahedral continuous image variant; For each horizontal boundary fusion image, the pixel column in the image is used as the reverse loop scrolling object. Each pixel is shifted in the reverse loop along the horizontal direction by -1 / 2 of the image width, so that each pixel column moves in the opposite direction to the horizontal loop scrolling relative to the current position of the horizontal boundary fusion image. For pixel columns that exceed the left or right boundary of the image after reverse cyclic displacement, they are cyclically written to the other side boundary of the image according to the corresponding displacement distance, so that the horizontal boundary merges all pixel columns in the image to complete the wrap-around rearrangement while keeping the color information and vertical position of each pixel unchanged. Through the reverse cyclic scrolling, the original left boundary and the original right boundary regions that were moved to the horizontal center of the image during the horizontal cyclic scrolling are moved back to their original coordinate positions. The image content after horizontal seam repair and boundary fusion processing is synchronously moved back to the original left and right boundary positions, generating a horizontal seamless image consistent with the original image coordinate positions. Output N seamless horizontal images after reverse looping processing, so that the leftmost and rightmost boundary pixel columns of each seamless horizontal image form a continuous transition relationship during loop stitching, thereby satisfying the seamless repeat stitching condition in the horizontal direction for each seamless horizontal image.
[0027] In this embodiment, step eight specifically includes: The N horizontal seamless images are cyclically scrolled along the vertical direction. The pixels in each horizontal seamless image are vertically cyclically displaced according to 1 / 2 of the image height, so that the original upper and lower boundaries of the horizontal seamless image are moved to the vertical center position of the image, generating N vertically cyclically scrolling images. A horizontal strip binary mask is generated based on the vertical center position of each vertically looping image. The area covered by the horizontal strip binary mask is determined as the area to be repaired for the vertical seam. The pixels in the area to be repaired for the vertical seam are conditionally masked to generate N vertical strip mask images. Scaling, ImageNet mean and standard deviation normalization, and zeroing of mask region pixels are performed on N vertical band mask images respectively to generate N vertical visible region images; N vertical visible region images are input into a mask autoencoder along the batch dimension. The intermediate features of the 6th layer decoder are extracted to generate vertical visual semantic features corresponding to each quad continuum image variant. The vertical visual semantic features are conditionally aligned by the CLIP alignment submodule and the T5 alignment submodule respectively to generate vertical CLIP condition vectors and vertical T5 condition sequences corresponding to each quad continuum image variant. N vertical strip mask images are batch compressed into the latent space using a variational autoencoder to generate vertical image latent codes. The horizontal strip binary mask is downsampled and rearranged. The processed mask data and the vertical image latent codes are concatenated along the channel dimension to generate vertical joint condition inputs. Independent Gaussian noise is initialized according to the random seeds corresponding to each quadruple continuous graph variant. N Gaussian noises and N vertical joint condition inputs are concatenated along the batch dimension. The concatenation result, the corresponding vertical CLIP condition vector, and the vertical T5 condition sequence are input into the rectified flow diffusion model. A second batch iterative denoising is performed with 50-step rectified flow time step scheduling, time shift scheduling, and a guiding scale of 30 to generate N vertical repair latent codes. The N vertical restoration latent codes are input into a standard variational autoencoder decoder for batch decoding to generate N vertical restoration RGB images. Gaussian pyramids and Laplacian pyramids are constructed based on each vertical restoration RGB image, the corresponding vertical cyclic scrolling image, and the horizontal bar binary mask. Laplacian layer weighted fusion is performed on each pyramid layer according to the corresponding layer mask, and the fused Laplacian pyramid is reconstructed layer by layer to generate N vertical boundary fused images. Each of the N vertical boundary fusion images is subjected to reverse cyclic scrolling along the vertical direction. The reverse cyclic scrolling distance is negative 1 / 2 of the image height, so that the vertical boundary fusion image is restored to the original image coordinate position, generating N square continuous pattern images. The left and right boundaries of each square continuous pattern image form a continuous stitching relationship, and the upper and lower boundaries form a continuous stitching relationship, outputting N square continuous images generated based on different random seeds.
[0028] refer to Figure 2 A novel four-square continuous graph generation system based on image inpainting includes the following modules: The horizontal mask processing module is used to acquire the RGB pattern image to be processed, allocate different random seeds according to the preset variant generation quantity, perform horizontal cyclic scrolling on the RGB pattern image, and generate a vertical strip binary mask to obtain a horizontal masked image. The visual semantic extraction module is used to extract visual features from the horizontal masked image through a mask autoencoder to generate horizontal visual semantic features. The dual-path condition alignment module is used to input the horizontal visual semantic features into the dual-path alignment submodule for feature transformation, and obtain horizontal visual condition information including the horizontal CLIP condition vector and the horizontal T5 condition sequence. The latent space coding module is used to input the horizontal band mask image into the variational autoencoder, compress the horizontal band mask image into the latent space, perform downsampling and rearrangement, and initialize Gaussian noise with a random seed to generate horizontal batch noise input; The rectified flow denoising module is used to input the horizontal batch noise input and the horizontal visual condition information into the rectified flow diffusion model for iterative denoising, generating N horizontal repair latent codes; The decoding and fusion module is used to batch decode the N horizontal repair latent coding inputs to the standard variational autoencoder decoder, and generate N horizontal boundary fusion images through Laplacian pyramid fusion and boundary smoothing processing. The horizontal seamless generation module is used to perform reverse cyclic rolling on N horizontal boundary fusion images respectively to generate N horizontal seamless images; The vertical seamless generation module is used to perform vertical cyclic scrolling on N horizontal seamless images respectively, repeating steps one to seven to obtain a four-sided continuous image.
[0029] Example 1: To verify the feasibility of this invention in practice, an example is given of its application to the continuous printing pattern production scenario of a textile pattern design center. This design center routinely needs to convert single RGB floral, foliage, and geometric texture patterns provided by designers into continuous four-dimensional patterns that can be repeatedly laid out horizontally and vertically. The original patterns provided by designers are usually independently completed single images, with no pre-established pattern connection relationship between their left and right edges and top and bottom edges. When the original patterns are directly repeated horizontally or vertically, petal truncation, misalignment of branches and leaves, abrupt color changes, and texture breaks easily occur at the edges. For images with dense pattern elements, the left and right sides may contain flower outlines of different scales, and direct splicing will create obvious vertical dividing lines at the boundaries; for images with continuous branch and leaf directions, the branch directions between the upper and lower boundaries are different, easily forming periodic horizontal seams when laid vertically. The original processing method at this design center required designers to first move the image boundary positions, then check the seams one by one, and repeatedly modify the seam areas. This typically resulted in a long processing time for a single complex pattern, and different designers produced different results for the same pattern. To address the issues of inconsistencies between the repaired content and the original image's visual features, obvious transitions in the repaired area boundaries, and the need to repeatedly generate multiple variations of complex RGB patterns, the four-sided continuous image generation method based on image inpainting described in this invention is applied in this design scenario.
[0030] In this embodiment, an RGB floral pattern is selected from the design center's pattern library as the image to be processed. The original image consists of dark green leaves, light green branches, red petals, and a light-colored background, with a pattern width of 2048 pixels and a pattern height of 2048 pixels. When the original image is repeatedly laid out for observation, there are three obvious petal truncated areas and two branches and leaves texture misalignment areas at the left and right boundaries, and four branches broken areas at the top and bottom boundaries. After directly repeating the original image in a 3×3 pattern, continuous vertical seams can be observed in the horizontal direction, and periodic horizontal texture abrupt changes can be observed in the vertical direction. The design center selects the need to obtain four four-dimensional continuous image variants simultaneously according to the actual scheme, so the preset variant generation quantity is set to 4, and the number of four-dimensional continuous image variants N is determined to be 4. The system assigns different random seeds to the four four-dimensional continuous image variants, so that the subsequent Gaussian noise initialization process has independent noise input, generating four different repair results while maintaining the same original RGB pattern content.
[0031] After reading the RGB pattern image, the system first determines the image width and height, and establishes the processing correspondence between the original image, the number of variants, and the random seed. Then, it performs horizontal cyclic scrolling on the original RGB pattern image. Since the image width is 2048 pixels, the horizontal cyclic displacement distance is half the image width, or 1024 pixels. Each pixel in the original image moves 1024 pixels horizontally. During the movement, pixels exceeding one boundary cycle into the other side, thus moving the pattern content corresponding to the left and right boundaries of the original image to the central area of the horizontally scrolling image. After cyclic scrolling, the red petal truncated areas, dark green leaf edges, and light green branch misalignment areas originally located on both sides of the image converge near the center of the image. The original external boundary problem is transferred to the internal seam repair problem of the image.
[0032] A vertical stripe binary mask is generated at the horizontal center of the horizontally scrolling image. The central area covered by this mask is used as the area to be repaired for the horizontal seam, thus obtaining a horizontal image with the mask. Actual observation shows that the original seam areas formed by the left and right boundaries are all within the coverage of the vertical stripe mask, while the complete flower, leaf, and background textures far from the seam remain outside the mask. This process avoids regenerating all pattern content of the entire 2048×2048 pixel image; instead, subsequent repair processing is focused on the central seam area formed after the scrolling process.
[0033] The horizontal masked image is scaled to 256×256 pixels and normalized according to the mean and standard deviation of the ImageNet dataset. After normalization, the pixel values of the corresponding areas of the vertical bar binary mask are cleared to zero, retaining only the visible pattern content outside the mask to generate a horizontal visible area image. This horizontal visible area image still retains the color distribution of red petals, the outline of dark green leaves, the direction of light green branches, and the overall visual relationship of the background area. The horizontal visible area image is input into a mask autoencoder, where visual features are extracted through the encoder and decoder, and the intermediate features of the 6th layer decoder are read. The intermediate features corresponding to each image patch are represented as 512-dimensional features, which are combined to form an N×512-dimensional horizontal visual semantic feature. Through this processing, the conditional information used for subsequent restoration comes from the current pattern to be processed itself, rather than additional text descriptions separate from the pattern content.
[0034] The horizontal visual semantic features are input into the CLIP alignment submodule and the T5 alignment submodule, respectively. The CLIP alignment submodule converts the horizontal visual semantic features into a horizontal CLIP condition vector with an output dimension of 768, and the T5 alignment submodule converts the horizontal visual semantic features into a horizontal T5 condition sequence with an output dimension of 4096. After processing, the horizontal CLIP condition vector is used as the first visual condition information, and the horizontal T5 condition sequence is used as the second visual condition information. The two are then combined to form the horizontal visual condition information. In the floral pattern of this embodiment, both conditions originate from the same horizontal visual semantic feature. Therefore, during the denoising process of the rectified flow diffusion model, the generation of the central seam region is constrained by the visual information of the visible area of the original image.
[0035] Simultaneously, the horizontal masked image input is processed using a variational autoencoder with float32 precision. The variational autoencoder compresses the horizontal masked image to a latent space with 16 channels and a spatial size downsampled by 8 times, forming the horizontal image latent code. The vertical strip binary mask is downsampled in 8×8 blocks and then rearranged using a 2×2 packing method. The processed mask data is then concatenated with the horizontal image latent code along the channel dimension, thus forming the horizontal joint conditional input. For the four quadruple continuous image variants, the system initializes independent Gaussian noise with the same shape as the horizontal image latent code using four random seeds. The four Gaussian noises and four sets of horizontal joint conditional inputs are then concatenated along the batch dimension to generate the horizontal batch noise input. Thus, the four variants maintain the same pattern conditions and different initial noise states during the same batch processing.
[0036] The horizontal batch noise input and horizontal visual condition information are input into the rectified flow diffusion model. The system generates a 50-step rectified flow time-step schedule and uses time-shift scheduling to adjust the denoising rhythm at the corresponding resolution. The guiding scale is set to 30, and the forward calculation is repeatedly performed through the diffusion transformer within the 50 denoising time steps. The horizontal CLIP condition vector and the horizontal T5 condition sequence jointly determine the denoising guiding direction corresponding to each time step. The latent codes of the four variants are updated separately in the same batch denoising process. After 50 iterations, the latent space content corresponding to the central horizontal seam to be repaired gradually forms a pattern feature that matches the original floral pattern. After denoising, each latent code is restored from the sequence form to the spatial form, obtaining four horizontal repair latent codes.
[0037] Four horizontal inpainting latent codes were input into a standard variational autoencoder decoder for batch decoding, outputting four horizontally inpainted RGB images. Upon examination of the four images, it is evident that new connections have been generated at the original central seam location. However, subtle differences exist in the results corresponding to different random seeds. Some variants form continuous petal edges at petal truncation points, while others exhibit different leaf connection patterns at branch-leaf junctions. Subsequently, an L-layer Gaussian pyramid was constructed based on the horizontally inpainted RGB images, the corresponding horizontally looping image, and a vertical bar binary mask. Gaussian downsampling was performed simultaneously on the vertical bar binary mask, and Laplacian pyramids for each image were generated based on adjacent Gaussian pyramid layers. Weighted fusion was performed on the Laplacian layers of the horizontally inpainted RGB images and the horizontally looping image at different pyramid layers based on the corresponding layer masks. The fused Laplacian pyramids were reconstructed layer by layer from coarse to fine, resulting in four horizontally boundary-fused images. After fusion, the color and texture transitions between the generated content in the central seam region and the surrounding original pattern are more continuous.
[0038] Four horizontally bounded images are subjected to a reverse loop scrolling along the horizontal direction. The reverse loop scrolling distance is half the image width, i.e., 1024 pixels in the opposite direction to the previous horizontal loop scrolling. The left and right boundary content that had previously moved to the center of the image returns to its original coordinate position. The repaired and blended seam content synchronously returns to the left and right edges of the image, generating four horizontally seamless images. After repeating the four horizontally seamless images horizontally, the three petal truncated areas in the original image no longer form obvious vertical seams, and the two misaligned areas of branch and leaf texture can be continuously connected at the left and right boundaries.
[0039] Vertical cyclic scrolling is then performed on the four horizontally seamless images. Since the image height is 2048 pixels, the vertical cyclic scrolling distance is 1024 pixels, moving the corresponding regions of the original upper and lower boundaries to the vertical center of the image. Subsequently, a corresponding mask is generated according to the aforementioned processing procedure, and the following steps are performed again: mask autoencoder visual feature extraction, CLIP alignment submodule and T5 alignment submodule dual-path conditional transformation, variational autoencoder latent space encoding, mask downsampling and rearrangement, independent Gaussian noise batch construction, 50-step rectified flow diffusion model iterative denoising, standard variational autoencoder batch decoding, and Laplacian pyramid fusion processing. Since the second processing object is now four different horizontally seamless images, each corresponding to its own pattern content, a corresponding reverse cyclic scrolling is performed after vertical restoration, ultimately outputting four quadruple continuous images. Observations are conducted by repeatedly laying 3×3 and 5×5 tiles on the four output images, and no obvious periodic seams as in the original pattern appear on the left and right boundaries or the top and bottom boundaries.
[0040] To verify the actual processing effect of the method of the present invention in this scenario, a continuous image generation test was conducted on 120 RGB pattern images on the image processing workstation of the textile pattern design center. The test images included 45 floral images, 31 leaf images, 24 geometric texture images, and 20 mixed pattern images, with image sizes of 1024×1024 pixels, 1536×1536 pixels, and 2048×2048 pixels. In the test, the results of existing manual cyclic scrolling followed by local repair processing were used as the comparison results, and the results generated by the method of the present invention were used as the example processing results. The number of abnormal seam images was counted by repeating the output image 3×3 to count the number of images with obvious left and right boundary or top and bottom boundary breaks; the average color difference of the boundary was the average value obtained by statistically analyzing the color difference of adjacent pixels on both sides of the splicing boundary; the average processing time per image was the average time required to complete an original pattern and obtain a usable four-sided continuous image; the number of output images per processing session represented the number of variant images obtained in one processing task. The test results are shown in the table below.
[0041] Table 1. Comparison of Processing Effects and Running Performance for Generating Complex RGB Patterns in a Four-Sided Continuous Pattern As shown in Table 1, in the comparative test of 120 pattern images, after manual cyclic scrolling and local repair processing, 21 images still had horizontal seam anomalies, and 25 images had vertical seam anomalies. After applying the method of this invention, these numbers decreased to 5 and 6 respectively. Images with both horizontal and vertical seam anomalies decreased from 12 to 2, indicating that by first performing horizontal cyclic scrolling and repair on the RGB pattern images, and then performing vertical corresponding processing on N seamless horizontal images, continuous processing can be achieved for both the left and right boundaries and the top and bottom boundaries. Under the 3×3 repeated laying condition, the number of images with obvious seams decreased from 33 to 7. After expanding the laying range to 5×5, the number of images with visible seams in the manual processing results increased to 38, while the method of this invention resulted in 9. This shows that the output image of this invention maintains a good boundary continuity even after the repeated range is expanded.
[0042] Based on the boundary color difference data, the average color difference between the left and right boundaries in the manually processed result was 8.74, and the average color difference between the top and bottom boundaries was 9.18. The corresponding values after processing by the method of this invention were 3.12 and 3.46, respectively. The number of images with obvious edges in the repaired area was reduced from 29 to 6. This is because after the standard variational autoencoder decoder completes batch decoding, it does not directly use the repaired image as the final output. Instead, it constructs Gaussian and Laplacian pyramids based on the horizontally repaired RGB image, the looping image, and the binary mask, respectively. Weighted fusion is performed at different pyramid layers based on the mask, and reconstruction is performed layer by layer from coarse to fine. Therefore, the repaired content and the original image content are not simply replaced at the boundary positions, which reduces obvious differences in regional edges.
[0043] Regarding pattern structure representation, the number of images with abrupt changes in petal or main outline during the original comparison processing was 18, which was reduced to 4 after using the method of this invention; the number of images with misaligned branch and leaf textures was reduced from 24 to 5. Observing specific flower pattern results, after the mask autoencoder extracts visual features from the horizontally visible area image, the intermediate features of the 6th layer decoder can form horizontal visual semantic features. These horizontal visual semantic features are converted into a 768-dimensional horizontal CLIP conditional vector and a 4096-dimensional horizontal T5 conditional sequence by the CLIP alignment submodule and T5 alignment submodule, respectively. These two visual conditions jointly participate in the 50-step iterative denoising of the rectified flow diffusion model, ensuring that the latent space restoration process is continuously constrained by the visual features of the original visible area. The resulting restored content exhibits good continuity with the original pattern in terms of petal color, leaf texture, and branch direction.
[0044] In terms of processing time, the average processing time per image for manual cyclic scrolling and local repair methods is 18.6 min, while the average processing time per image for the method of this invention is 2.7 min. For 2048×2048 pixel images, the average processing time for manual methods is 24.3 min, while the average processing time for the method of this invention is 3.4 min. When four optional pattern variants are required, the manual method processes and calculates them separately, resulting in a total average generation time of 74.4 min for the four variants. In contrast, this invention initializes independent Gaussian noise with different random seeds and stitches the four Gaussian noises and the corresponding four joint condition inputs along the batch dimension, performing batch latent space repair in the rectified flow diffusion model. The total average generation time for the four variants is 8.9 min. The cumulative processing time for 120 test images is reduced from 37.2 h to 5.4 h, which meets the application requirements of the design center for continuously processing multiple pattern materials.
[0045] In this implementation scenario, the present invention does not simply move the boundaries of the original image, but rather moves the horizontal and vertical seams sequentially into the image. It extracts the visual semantic features of the current pattern itself through a mask autoencoder, and generates horizontal visual conditional information through the CLIP alignment submodule and T5 alignment submodule. Then, a rectified flow diffusion model is used to complete latent space inpainting guided by visual conditions. Standard variational autoencoder decoding and Laplacian pyramid fusion are used to reduce edge differences in the inpainting. The combination of different random seeds and batch noise input allows the same original RGB pattern to generate multiple different tetrahedral continuous image variants in a single processing task. Test results show that this method can reduce horizontal and vertical seams, color abrupt changes, and local texture misalignment when repeatedly laying complex RGB patterns, while shortening the generation time of multiple pattern variants, demonstrating the practical application effect of the present invention in the automatic generation of tetrahedral continuous images.
[0046] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A novel method for generating four-sided continuous graphs based on image inpainting, characterized in that, include: Step 1: Obtain the RGB pattern image to be processed, assign different random seeds according to the preset variant generation quantity, perform horizontal cyclic scrolling on the RGB pattern image, and generate a vertical strip binary mask to obtain a horizontal masked image; Step 2: Extract visual features from the horizontal masked image using a mask autoencoder to generate horizontal visual semantic features; Step 3: Input the horizontal visual semantic features into the dual-path alignment submodule for feature transformation to obtain horizontal visual condition information including the horizontal CLIP condition vector and the horizontal T5 condition sequence; Step 4: Input the horizontal masked image into the variational autoencoder to compress the horizontal masked image into the latent space, perform downsampling and rearrangement, and initialize Gaussian noise with a random seed to generate horizontal batch noise input; Step 5: Input the horizontal batch noise input and horizontal visual condition information into the rectified flow diffusion model for iterative denoising to generate N horizontal repair latent codes; Step 6: Input the N horizontal insulated latent codes into the standard variational autoencoder decoder for batch decoding, and generate N horizontal boundary fusion images through Laplacian pyramid fusion and boundary smoothing processing; Step 7: Perform reverse loop rolling on the N horizontal boundary fusion images respectively to generate N horizontal seamless images; Step 8: Perform vertical cyclic scrolling on the N horizontal seamless images respectively, repeating steps one to seven to obtain a four-sided continuous image.
2. The novel method for generating a tetrahedral continuous graph based on image inpainting according to claim 1, characterized in that, Step one specifically involves: Obtain the RGB pattern image to be processed, determine the image width and image height of the RGB pattern image, and determine the number N of tetrahedral continuous image variants to be generated according to the preset number of variants generated. Assign different random seeds to each tetrahedral continuous image variant to generate the pattern image processing task. For each pattern image processing task, the RGB pattern image is cyclically scrolled horizontally, and each pixel in the image is cyclically shifted horizontally by 1 / 2 of the image width to generate a horizontally scrolling image. A vertical strip binary mask is generated based on the horizontal center position of the horizontally looping image. The area covered by the vertical strip binary mask is determined as the horizontal seam to be repaired, and a horizontal masked image is generated.
3. A novel method for generating a tetrahedral continuous graph based on image inpainting according to claim 1, characterized in that, Step two specifically involves: The horizontal masked image is scaled to 256×256 pixels and normalized according to the mean and standard deviation of the ImageNet dataset. Clear the pixel values of the region corresponding to the vertical bar binary mask in the normalized image to zero, and generate the horizontally visible region image; The horizontally visible area image is input into a mask autoencoder. Visual features are extracted through the encoder and decoder of the mask autoencoder. The intermediate features of the 6th layer decoder of the mask autoencoder are extracted to generate horizontal visual semantic features with a dimension of N×512.
4. A novel method for generating a tetrahedral continuous graph based on image inpainting according to claim 1, characterized in that, Step three specifically involves: The dual-path alignment submodule includes a CLIP alignment submodule and a T5 alignment submodule; The horizontal visual semantic features are input into the CLIP alignment submodule and the T5 alignment submodule, respectively. The CLIP alignment submodule converts the horizontal visual semantic features into a horizontal CLIP conditional vector with an output dimension of 768, and the T5 alignment submodule converts the horizontal visual semantic features into a horizontal T5 conditional sequence with an output dimension of 4096. The horizontal CLIP condition vector is used as the first visual condition information, the horizontal T5 condition sequence is used as the second visual condition information, and the first and second visual condition information are combined into the horizontal visual condition information.
5. A novel method for generating a tetrahedral continuous graph based on image inpainting according to claim 1, characterized in that, Step four specifically involves: The horizontal masked image is input into a variational autoencoder with float32 precision, and the horizontal masked image is compressed to a latent space with 16 channels and a spatial size downsampled by 8 times to generate a horizontal image latent code. The vertical strip binary mask is downsampled in an 8×8 block manner and rearranged in a 2×2 packing manner. The processed mask data is then spliced with the horizontal image latent code along the channel dimension to generate a horizontal joint conditional input. Each independent Gaussian noise with the same latent coding shape as the horizontal image is initialized according to the random seed corresponding to each four-dimensional continuous graph variant. The N Gaussian noises and the corresponding N horizontal joint conditional inputs are then concatenated along the batch dimension to generate the horizontal batch noise input.
6. A novel method for generating a tetrahedral continuous graph based on image inpainting according to claim 1, characterized in that, Step five specifically involves: The horizontal batch noise input and horizontal visual condition information are input into the rectified flow diffusion model for iterative denoising to generate a 50-step rectified flow time step schedule. Time-shift scheduling is used to adjust the denoising rhythm for images of different resolutions; A 50-step iterative denoising process is performed with a guiding scale of 30. Forward calculation is performed through a diffusion transformer at each denoising time step. The denoising guiding direction is determined by the horizontal CLIP condition vector and the horizontal T5 condition sequence. Batch latent space repair is performed on the horizontal seam repair areas corresponding to N quadrature continuous graph variants. The denoised latent codes are restored from sequence form to spatial form, generating N horizontally repaired latent codes.
7. A novel method for generating a four-sided continuous graph based on image inpainting according to claim 1, characterized in that, Step six specifically involves: The N horizontal restoration latent codes are input into a standard variational autoencoder decoder for batch decoding to generate N horizontal restoration RGB images; An L-layer Gaussian pyramid is constructed based on each horizontally repaired RGB image, the corresponding horizontally looping image, and the vertical strip binary mask. The vertical strip binary mask is simultaneously downsampled by Gaussian, and the Laplacian pyramids corresponding to the horizontally repaired RGB image and the horizontally looping image are generated based on the adjacent Gaussian pyramid layers. In each pyramid layer, the Laplacian layer of the horizontally repaired RGB image and the Laplacian layer of the horizontally looping image are weighted and fused according to the corresponding layer mask. The fused Laplacian pyramid is then reconstructed layer by layer from coarse to fine to generate N horizontal boundary fused images.
8. A novel method for generating a tetrahedral continuous graph based on image inpainting according to claim 1, characterized in that, Step seven specifically involves: Perform reverse loop scrolling along the horizontal direction on each of the N horizontal boundary fused images; The reverse loop rolling distance is negative 1 / 2 of the image width, until the horizontal boundary fused image is restored to the original image coordinate position, generating N horizontal seamless images; The left and right boundaries of seamless images at each level form a continuous stitching relationship.
9. A novel four-sided continuous graph generation system based on image inpainting, comprising the novel four-sided continuous graph generation method based on image inpainting as described in any one of claims 1 to 8, characterized in that, Includes the following modules: The horizontal mask processing module is used to acquire the RGB pattern image to be processed, allocate different random seeds according to the preset variant generation quantity, perform horizontal cyclic scrolling on the RGB pattern image, and generate a vertical strip binary mask to obtain a horizontal masked image. The visual semantic extraction module is used to extract visual features from the horizontal masked image through a mask autoencoder to generate horizontal visual semantic features. The dual-path condition alignment module is used to input the horizontal visual semantic features into the dual-path alignment submodule for feature transformation, and obtain horizontal visual condition information including the horizontal CLIP condition vector and the horizontal T5 condition sequence. The latent space coding module is used to input the horizontal band mask image into the variational autoencoder, compress the horizontal band mask image into the latent space, perform downsampling and rearrangement, and initialize Gaussian noise with a random seed to generate horizontal batch noise input. The rectified flow denoising module is used to input the horizontal batch noise input and the horizontal visual condition information into the rectified flow diffusion model for iterative denoising, generating N horizontal repair latent codes; The decoding and fusion module is used to batch decode the N horizontal repair latent coding inputs to the standard variational autoencoder decoder, and generate N horizontal boundary fusion images through Laplacian pyramid fusion and boundary smoothing processing. The horizontal seamless generation module is used to perform reverse cyclic rolling on N horizontal boundary fusion images respectively to generate N horizontal seamless images; The vertical seamless generation module is used to perform vertical cyclic scrolling on N horizontal seamless images respectively, repeating steps one to seven to obtain a four-sided continuous image.