AIGC image rapid generation method and system

By using image block segmentation and multi-scale texture compression feature extraction, combined with semantic residual mapping and quality assessment, the consistency and robustness issues of image generation in existing technologies are solved, achieving efficient and stable image generation results.

CN120976351AInactive Publication Date: 2025-11-18ZHEJIANG QINGDA TECH IND CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511153030.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-11-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing image generation methods fail to effectively distinguish the complexity of local image content, resulting in unstable quality of local detail reconstruction. They lack detailed analysis of the structural consistency and semantic continuity of local image regions and lack a unified quality assessment and feedback adjustment mechanism, which affects the consistency and robustness of generated images.

Method used

By dividing the image into blocks, extracting multi-scale texture compression features, semantic residual mapping, and evaluating image quality, a compressed texture index map and a semantic residual mapping map are constructed. A multi-channel generation model is used to generate images, and the generation strategy is adjusted iteratively through feedback optimization.

Benefits of technology

It achieves a fine response to the structural complexity of image regions, alleviates the semantic mismatch and structural breakage problems in local generation, improves the stability and control precision of generated images, and enhances generation efficiency and task adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976351A_ABST
    Figure CN120976351A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image data processing, and discloses an AIGC image rapid generation method and system, and the method comprises the steps: obtaining an original image source, dividing the original image source into image blocks, constructing an image block set based on a multilayer pyramid rule, extracting the frequency domain sparseness, the space gradient density, the wavelet energy distribution and other texture features, and carrying out the rapid generation of an AIGC image. Generating, sorting and compressing a texture index map; analyzing the semantic structure of the image block, constructing a semantic background image, calculating a semantic residual error, generating a semantic residual error mapping image, and inputting the semantic residual error mapping image into an AIGC model to generate the content of the image block; fusing the plurality of image blocks to generate a preliminary image, evaluating image quality and extracting texture, color and structure indexes; and optimizing a generation path based on an evaluation result, realizing iterative updating of image reconstruction, and outputting a final image. According to the method, by constructing a dual guide mechanism of the compressed texture index map and the semantic residual mapping map, the control of the AIGC generated image on the structure and semantic level is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to an AIGC image rapid generation method and system. BACKGROUND

[0002] With the rapid development of artificial intelligence generated content (AIGC, i.e. Artificial Intelligence Generated Content) technology, its application in the field of image generation is continuously expanding, and it has been widely used in image restoration, blur removal, structure completion and content enhancement tasks. The existing image generation method generally relies on deep learning model to process the whole image in an end-to-end manner. Typical methods such as image restoration model based on semantic segmentation map or feature point can restore the structure and color information of the image to a certain extent.

[0003] However, the existing technology has the following problems: first, most of the generation processes do not distinguish the local content complexity of the image, and cannot control the difference according to the texture density, structure change and other characteristics of different regions, resulting in unstable local detail reconstruction quality; second, semantic information is often used only for overall condition constraints, lacking fine analysis of image local region structure consistency, semantic continuity and other aspects, which is prone to problems such as boundary jump and semantic fracture of generated image; in addition, the image generation process generally uses one-way inference mechanism, lacks unified quality evaluation and feedback adjustment mechanism, and cannot automatically correct the previous processing strategy according to the generation result, affecting the consistency and robustness of the final image output. SUMMARY

[0004] In view of this, the present application provides an AIGC image rapid generation method and system to solve the above problems.

[0005] In one aspect, the present application provides an AIGC image rapid generation method, comprising: obtaining an original image source, dividing the original image source into image blocks, constructing an image block set based on a multi-layer pyramid rule, and extracting multi-scale texture compression features of each image block, including frequency domain sparsity, spatial gradient density and wavelet energy distribution; constructing a compressed texture index atlas based on the multi-scale texture compression features, sorting according to the reconstruction priority of the image block, and taking the compressed texture index atlas as the guide basis of the AIGC generation path; performing semantic structure analysis on each image block region, constructing a corresponding regional semantic background map, and performing semantic residual calculation on the original image block and the regional semantic background map to generate a semantic residual mapping map; inputting the semantic residual mapping map into a preset AIGC generation model to complete the generation of the target image block content; The plurality of generated image blocks are subjected to boundary fusion processing, and an edge jump detection and color histogram matching strategy is adopted to perform dynamic fusion of edge regions to generate a preliminary synthesized image. The preliminary synthesized image is subjected to image quality evaluation, and image generation stability indicators including texture continuity, color consistency and structural integrity are extracted. The image generation stability indicator results are fed back to the construction process of the compressed texture index map and the semantic residual mapping map, the reconstruction priority ranking and residual information extraction strategy are updated, the iteration optimization processing of the image reconstruction path is performed, and the final image generation result is output.

[0006] Further, when the original image source is subjected to image block division based on the multi-layer pyramid rule, it includes: After obtaining the original image source, the original image is initially divided according to a fixed size to obtain a plurality of image blocks; Each image block is subjected to the following operations: The high-frequency sub-band of the image block is extracted using discrete wavelet transform, and the energy density value in the high-frequency sub-band is calculated to represent the texture complexity; The Sobel edge detection is performed on the image block, the gradient direction of all edge pixels is counted, the ratio of the number of main directions to the number of other directions is calculated, and the edge direction consistency score is obtained; The mean and standard deviation of the RGB three color channels of the image block are calculated, and the variance superposition value between color channels is calculated to represent the color composition difference; The texture complexity, edge direction consistency score and color composition difference are input as feature vectors into the pre-set multi-layer pyramid mapping rule, and the mapping rule matches the feature vectors in the multi-dimensional feature space with the reference threshold interval of each pyramid level to determine the pyramid level to which the image block belongs; The image blocks belonging to the same pyramid level in the determination result are grouped together, and a multi-layer pyramid structure including a plurality of level image block sets is finally constructed.

[0007] Further, when extracting the multi-scale texture compression features of each image block, it includes: For each image block, three feature extraction paths are used respectively: (1) Frequency domain sparsity extraction path: the image block is subjected to discrete Fourier transform to obtain frequency domain coefficient distribution, the proportion of frequency domain coefficients with amplitude greater than a preset amplitude is counted, and is recorded as frequency domain sparsity; (2) Spatial gradient density extraction path: the Sobel operator is used to extract the gradient image of the image block in the horizontal and vertical directions, the gradient image is subjected to pixel counting, the proportion of pixels with gradient amplitude exceeding the set gradient threshold is counted, and the gradient direction variance is calculated to form the spatial gradient density; (3) Wavelet energy distribution extraction path: based on Daubechies wavelet, three-layer wavelet decomposition is performed on the image block, LL, LH, HL and HH subbands are obtained, the energy of each subband is normalized, the energy distribution vector is constructed, and the wavelet energy distribution features are represented by energy concentration and subband energy ratio; The frequency domain sparsity, spatial gradient density and wavelet energy distribution are combined into a compressed texture feature vector as the basis index input for the priority ranking of the construction and reconstruction of the compressed texture index atlas.

[0008] Further, when constructing the compressed texture index atlas, it includes: A three-dimensional feature mapping space is established, and the frequency domain sparsity, spatial gradient density and wavelet energy distribution of the image block are respectively mapped to the three coordinate axes to form a compressed texture feature scatter plot; In the compressed texture feature scatter plot, a local density clustering algorithm is used to identify high-density feature clusters, and by setting the feature density mean and the radius range of the clustering center neighborhood, image blocks located in the high-density clustering and with three-dimensional coordinate differences within the preset tolerance threshold are marked to construct a compressed texture redundant block set; For image blocks that are not marked as compressed texture redundant blocks, the Euclidean distance between the coordinate points corresponding to the frequency domain sparsity, spatial gradient density and wavelet energy distribution in the three-dimensional feature mapping space and the feature distribution boundary is calculated, and the reconstruction priority score is generated according to the Euclidean distance and the spatial position boundary of the image block in the original image source; All image blocks are sorted in descending order of reconstruction priority score to construct a compressed texture index atlas, which includes the spatial coordinates, frequency domain sparsity, spatial gradient density, wavelet energy distribution and corresponding reconstruction priority score of the image block, and is used as the basis for guiding the AIGC generation path.

[0009] Further, when generating the semantic residual mapping map, it includes: Based on the pyramid level to which the image block belongs, a local image region with an area of one-tenth of the original image area is extracted in the original image source as the basis input of the regional semantic background map; The image block and the regional semantic background map are respectively input into the same semantic analysis model to perform semantic segmentation and semantic boundary extraction operations to extract the semantic label matrix, boundary contour map and structure level annotation map; The semantic label matrices of the image block and the regional semantic background map are compared to generate a semantic coverage difference map; The structure contour comparison is performed on the boundary contour maps of the image block and the regional semantic background map to generate a boundary offset map based on the boundary direction similarity, interface length and closure degree indexes; Perform hierarchical matching on the image block and the structure hierarchical label map of the regional semantic background map, count the hierarchical difference and nesting error, and generate a hierarchical deviation map; Register and superimpose the semantic coverage difference map, the boundary offset map, and the hierarchical deviation map in the image coordinate position to generate a semantic residual mapping map.

[0010] Further, the step of performing semantic structure analysis on each image block region comprises: Perform semantic segmentation on the image block using an image segmentation network to obtain a corresponding semantic label matrix; Extract a boundary contour map according to the semantic label matrix, and align the edges in combination with the grayscale image of the image block; Label the bounding rectangle frame of each semantic region in the semantic label matrix, and count the boundary contact length and semantic class combination mode between adjacent semantic regions; Construct a structure hierarchical label map based on the boundary contact relationship, which is used to represent the hierarchical nesting relationship and local structure topology of each semantic region in the image block.

[0011] Further, the step of performing semantic residual calculation on the original image block and the regional semantic background map comprises: Encode the original image block and the regional semantic background map into vector representations respectively, and extract semantic class feature distribution using a word embedding module; Based on the semantic nesting hierarchy recorded in the structure hierarchical label map, segment and pair the two vector representations; Perform cosine similarity calculation on each paired segment to construct a multi-segment semantic matching matrix; Extract the unaligned segment sequence in the multi-segment semantic matching matrix, and label it as a semantic residual segment through structure consistency analysis and color texture mutation recognition; Reconstruct all semantic residual segments into a semantic residual mapping map, which contains semantic segment position, semantic class, structure disconnection factor, and residual confidence score.

[0012] Further, the step of inputting the semantic residual mapping map into the preset AIGC generation model to complete the image block content generation comprises: Construct a prompt information input tensor based on the semantic residual mapping map, which contains the spatial coordinates, semantic class code, residual confidence score, and structure disconnection factor of each semantic residual segment; According to the pyramid level to which the image block belongs, call the corresponding generation sub-model, and input the original image block and the prompt information input tensor simultaneously using a multi-channel input method; Introduce an attention routing mechanism in the AIGC generation model, and adjust the activation weight of the multi-layer attention mechanism according to the semantic class code and residual confidence score in the prompt information input tensor; After generating the output image block, the output image block is compared with the original image block for local coincidence, and whether the semantic recovery degree and the structural continuity meet the standards is verified, and if not, the semantic residual error mapping is re-inputted for generation and update.

[0013] Further, the step of image quality evaluation comprises: The preliminary synthesized image is compared with the original image source at the image block level, and the texture continuity score, the color consistency score and the structural integrity score are calculated; The texture continuity score is obtained by counting the normalized value of the frequency response difference in the edge area of adjacent image blocks; The color consistency score is obtained by comparing the Bhattacharyya distance of the RGB color channel histogram of the boundary neighborhood of each image block; The structural integrity score is evaluated by checking whether the spatial connectivity and nested structure of the region semantic label matrix in the reconstructed image are consistent; The three types of scores form an image generation stability index group, which is used as the basis for iterative optimization of the final image generation result.

[0014] Compared with the prior art, the beneficial effects of the present application are: The present application extracts multi-dimensional texture features such as frequency domain sparsity, spatial gradient density and wavelet energy distribution, establishes a three-dimensional feature mapping space and generates a compressed texture index atlas, which not only realizes the hierarchical classification and reconstruction priority ordering of image blocks, but also improves the response capability to the structural complexity of image regions, ensuring that the reconstruction order is more targeted and global.

[0015] The present application constructs a region semantic background map and calculates semantic residual error to generate a residual error mapping containing semantic deviation, boundary misplacement and hierarchical conflict information, which assists the AIGC model in performing directed guidance generation, thereby effectively alleviating the problems of semantic mismatch and structural fracture in local generation.

[0016] The present application evaluates three types of indexes, namely texture continuity, color consistency and structural integrity, establishes an image quality evaluation feedback path, and applies the evaluation results to the dynamic adjustment process of the compressed texture index atlas and the semantic residual error strategy, realizes the iterative optimization of the generation strategy, and effectively improves the stability and control accuracy of the final generated image.

[0017] The present application modularly integrates image block division, feature extraction, index atlas construction, semantic residual error generation, image content generation, quality evaluation and feedback optimization, and constructs a unified linked image generation system, which avoids the problem of independent modules and interrupted adjustment path in the prior art, and improves the overall generation efficiency and task adaptability.

[0018] In another aspect, the application provides an AIGC image fast generation system, comprising: An image block construction module configured to obtain an original image source, perform image block division on the original image source, construct an image block set based on a multi-layer pyramid rule, and extract multi-scale texture compression features of each image block, including frequency domain sparsity, spatial gradient density, and wavelet energy distribution; A texture index atlas construction module configured to construct a compressed texture index atlas based on the multi-scale texture compression features, sort according to the reconstruction priority of the image block, and use the compressed texture index atlas as a guide for the AIGC generation path; A semantic residual mapping module configured to perform semantic structure analysis on each image block region, construct a corresponding regional semantic background map, and perform semantic residual calculation on the original image block and the regional semantic background map to generate a semantic residual mapping map; An image generation module configured to input the semantic residual mapping map into a preset AIGC generation model to complete the generation of the target image block content; A boundary fusion module configured to perform boundary fusion processing on multiple generated image blocks, use edge jump detection and color histogram matching strategies to perform dynamic fusion of edge regions, and generate a preliminary synthesized image; A quality evaluation module configured to perform image quality evaluation on the preliminary synthesized image, and extract image generation stability indicators including texture continuity, color consistency, and structural integrity; A feedback optimization module configured to feed back the image generation stability indicator results to the construction process of the compressed texture index atlas and the semantic residual mapping map, update the reconstruction priority sorting and residual information extraction strategy, perform iterative optimization processing of the image reconstruction path, and output the final image generation result.

[0019] It should be noted that the AIGC image fast generation method and system of the application have the same beneficial effects, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0020] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments. The drawings are for purposes of illustration only and are not considered a limitation of the application. Moreover, like reference numerals are used to designate identical components throughout the specification. In the drawings: Figure 1 A flowchart of an AIGC image fast generation method provided by an embodiment of the application.

[0021] Figure 2 A functional block diagram of an AIGC image fast generation system provided by an embodiment of the application. DETAILED DESCRIPTION

[0022] Exemplary embodiments of the present disclosure will be described in detail with reference to the drawings. Although exemplary embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure can be more thoroughly understood, and the scope of the present disclosure can be accurately conveyed to those skilled in the art. It should be noted that the embodiments in the present disclosure and the features in the embodiments can be combined with each other without conflict. The present disclosure will be described in detail below with reference to the drawings and in conjunction with the embodiments.

[0023] Reference Figure 1 As shown in the drawings, the embodiments of the present disclosure provide an AIGC image fast generation method, comprising: S1: obtaining an original image source, performing image block division on the original image source, constructing an image block set based on a multi-layer pyramid rule, and extracting multi-scale texture compression features of each image block, including frequency domain sparsity, spatial gradient density, and wavelet energy distribution; S2: constructing a compressed texture index atlas based on the multi-scale texture compression features, sorting according to the reconstruction priority of the image block, and taking the compressed texture index atlas as the guidance basis of the AIGC generation path; S3: performing semantic structure analysis on each image block region, constructing a corresponding regional semantic background map, and performing semantic residual calculation on the original image block and the regional semantic background map to generate a semantic residual mapping map; S4: inputting the semantic residual mapping map into a preset AIGC generation model to complete the generation of the target image block content; S5: performing boundary fusion processing on the plurality of generated image blocks, using edge jump detection and color histogram matching strategy to perform dynamic fusion of the edge region, and generating a preliminary synthesized image; S6: performing image quality evaluation on the preliminary synthesized image, and extracting image generation stability indicators including texture continuity, color consistency, and structural integrity; S7: feeding back the image generation stability indicator results to the construction process of the compressed texture index atlas and the semantic residual mapping map, updating the reconstruction priority sorting and residual information extraction strategy, performing iterative optimization processing on the image reconstruction path, and outputting the final image generation result.

[0024] In the embodiment, first, an original image source is acquired, which can be a static picture collected by an image collection device in real time or a high-resolution image file input in advance. In order to improve the spatial distribution control ability and detail preservation ability in the subsequent image generation process, the original image source is divided into a plurality of small-size image blocks in the embodiment. In the image block division process, the original image is uniformly divided by using a fixed-size sliding window mode, and the sliding window size and the overlap step length can be preset according to the image source resolution and the expected detail level, so as to ensure that the image content is completely covered and the image blocks have a certain context redundancy.

[0025] Further, in order to comprehensively describe the compression expression features of the image blocks in the frequency domain, spatial texture and structural level from the multi-scale perspective, the image block set is constructed based on the multi-layer pyramid rule in the embodiment. Specifically, according to the texture complexity, edge direction consistency and color difference degree of the image blocks, the divided image blocks are classified into different pyramid levels, so that the high-texture-complexity areas are included in the high layers of the pyramid, and the low-texture areas are included in the bottom layer, thereby forming a hierarchical image block distribution structure.

[0026] After the multi-layer pyramid structure is constructed, the multi-scale texture compression features of each image block are extracted, which are used to represent the local structure and compression complexity of the image content. The multi-scale features include the following three aspects: (1) frequency domain sparsity: the energy distribution sparsity of the image block in the frequency domain is calculated by using discrete Fourier transform; (2) spatial gradient density: the directional gradient detection is performed on the image block by using the Sobel operator, and the density distribution of the strong gradient response is counted; (3) wavelet energy distribution: the image block is decomposed by using multi-layer wavelet, and the energy concentration features of each frequency band are extracted. The three types of features can jointly constitute the compression texture feature vector of the image block, which is used as the basic index of the subsequent guided generation path.

[0027] In some embodiments of the present application, when the original image source is divided into image blocks and the image block set is constructed based on the multi-layer pyramid rule, the following steps are included: After the original image source is acquired, the original image is initially divided according to a fixed size, and a plurality of image blocks are obtained; The following operations are respectively performed on each image block: The high-frequency subband of the image block is extracted by using discrete wavelet transform, and the energy density value in the high-frequency subband is calculated to represent the texture complexity; The Sobel edge detection is performed on the image block, the gradient directions of all edge pixels are counted, the ratio of the number of main directions to the number of other directions is calculated, and the edge direction consistency score is obtained; The mean value and the standard deviation of the RGB three color channels of the image block are respectively calculated, the variance superposition value between the color channels is calculated, and the color composition difference degree is represented; The texture complexity, edge direction consistency score and color composition difference degree are input into a preset multi-layer pyramid mapping rule as a feature vector, the mapping rule matches the feature vector with reference threshold intervals of each pyramid level in a multi-dimensional feature space, and determines a pyramid level to which the image block belongs; Image blocks belonging to the same pyramid level in the determination result are grouped together, and a multi-layer pyramid structure including a plurality of level image block sets is finally constructed.

[0028] In the embodiment, in order to realize multi-scale structured expression of the image source content and provide a fine area division basis for subsequent texture feature extraction, compression sorting and semantic residual calculation, a combination strategy of image block division and multi-layer pyramid structure construction is adopted.

[0029] Firstly, after obtaining the original image source, the image is traversed and divided according to the set fixed size and sliding step. Specifically, a sampling window is slid in the horizontal and vertical directions of the image, and the original image is cut to obtain a set of image blocks with consistent size and covering the whole image. The image blocks can be non-overlapping or partially overlapping to ensure that each part of the image is mapped to at least one image block. Subsequently, in order to realize multi-dimensional texture expression and classification of the image blocks, each image block performs the following three feature extraction operations: texture complexity extraction, first, wavelet transform is performed on the image block to extract high-frequency subbands representing local changes. The energy density of the image block in the wavelet domain is calculated by counting the sum of squares of the coefficient values in these subbands, which is used to represent the richness of the detailed texture in the image block. The higher the energy density, the more dramatic the image content changes, and the higher the texture complexity.

[0030] Edge direction consistency score calculation, then, the image block is processed using an edge detection operator to identify all edge pixels and analyze the main gradient direction of these pixels. By comparing the frequency of the main direction with the frequency of other directions, the consistency of the edge direction is evaluated. If a certain direction dominates most of the edges, it means that the image block has strong structural directionality and clear outline, and the consistency score is high; if the edge direction distribution is scattered, the consistency score is low.

[0031] Color composition difference extraction, the pixel mean and fluctuation degree (i.e. standard deviation) of the image block in the red, green and blue color channels are calculated respectively, and then the difference between different color channels is counted. If the color distribution of the three channels is close, it means that the color composition is stable, and the difference degree is low; if there is a significant separation phenomenon between the color channels, the difference degree is high, which can be regarded as an image block with strong color diversity.

[0032] The three feature results are standardized and combined into a three-dimensional feature vector to comprehensively describe the performance of the image block in three dimensions of frequency domain, spatial structure, and color distribution.

[0033] Then, according to a preset pyramid level division rule, the feature vector of each image block is classified and judged. The rule sets a group of reference standards according to experience to divide three levels of “high complexity”, “medium complexity”, and “low complexity”. For example, when the texture complexity index of an image block is significantly higher than the preset standard, and the edge direction consistency is low, and the color composition difference is significant, it can be divided into the highest level; if the three indexes are in the middle interval, it is classified as medium level; if the texture complexity and color difference are both low, and the edge direction is highly consistent, it is classified as the lowest level.

[0034] According to the above judgment, image blocks with the same feature performance are classified into the same level, and finally a set of image blocks containing multiple levels is formed, and a complete multi-layer pyramid structure is constructed. The structure can effectively express the detail density and content difference of the image in the local area, and provide hierarchical support for the construction of the compressed texture index map in the subsequent.

[0035] In some embodiments of the present application, when extracting the multi-scale texture compression features of each image block, the following steps are included: For each image block, three feature extraction paths are used respectively: (1) Frequency domain sparsity extraction path: the image block is subjected to discrete Fourier transform to obtain the frequency domain coefficient distribution, the proportion of frequency domain coefficients with an amplitude greater than a preset amplitude is counted, and the frequency domain sparsity is recorded; (2) Spatial gradient density extraction path: the Sobel operator is used to extract the gradient image of the image block in the horizontal and vertical directions, the pixel count of the gradient image is performed, the proportion of pixels with a gradient amplitude exceeding a set gradient threshold is counted, and the gradient direction variance is calculated to form the spatial gradient density; (3) Wavelet energy distribution extraction path: based on the Daubechies wavelet, three-layer wavelet decomposition is performed on the image block to obtain LL, LH, HL, and HH subbands, the energy of each subband is normalized, an energy distribution vector is constructed, and the energy concentration and subband energy ratio are used to represent the wavelet energy distribution feature; The frequency domain sparsity, spatial gradient density, and wavelet energy distribution are combined into a compressed texture feature vector, which is used as the basic index input for the construction and reconstruction priority sorting of the compressed texture index map.

[0036] In the present fact example, in order to further improve the expression ability of the image block in the dimensions of content structure, texture details and color distribution, and to realize fine control of the image generation path, a multi-path fusion image block texture compression feature extraction strategy is proposed. Specifically, for each image block to be processed, the following three feature extraction paths are used for analysis and encoding: (1) Frequency domain sparsity extraction path: First, the discrete Fourier transform is performed on the image block to map it from the spatial domain to the frequency domain, obtaining the frequency domain coefficient distribution map. The frequency domain coefficients reflect the energy distribution of different frequency components in the image block. On this basis, a fixed amplitude threshold is selected, and the number of coefficients with an amplitude greater than the threshold is counted and compared with the total number of coefficients to obtain the frequency domain sparsity index. This index is used to measure the energy concentration of the image block in the frequency domain.

[0037] If the image block content is simple, there are only a few high-amplitude components in the frequency domain distribution, and the rest are near-zero values, then the frequency domain sparsity is high; on the contrary, if the high-amplitude frequency distribution is widespread, then the sparsity is low. This index can effectively describe the repetitive texture and compression potential of the image block.

[0038] (2) Spatial gradient density extraction path: The Sobel operator is used to perform convolution operations on the image block in the horizontal and vertical directions, respectively, to obtain the corresponding gradient images, which reflect the edge intensity changes of the image block in each direction. The gradient amplitude of each pixel in the gradient image is counted, and compared with the preset gradient threshold value to calculate the proportion of pixels greater than the threshold to the total number of pixels, which is used as the gradient response density index.

[0039] At the same time, the direction angle of all gradient pixels is counted, and the statistical variance value of the direction angle is calculated as the dispersion index of the direction distribution. The two indexes are combined to form a complete spatial gradient density feature. This feature is used to describe the concentration or dispersion of the number and direction of edges in the image block.

[0040] If an image block contains a large number of random edges, the gradient amplitude appears frequently and the direction is dispersed, indicating that the image structure is complex and the spatial gradient density is high; on the contrary, the gradient is sparse.

[0041] (3) Wavelet energy distribution extraction path: Based on Daubechies series wavelet (such as db4), the image block is processed by three-layer wavelet decomposition to obtain the coefficient matrix of LL (low frequency-low frequency), LH (low frequency-high frequency), HL (high frequency-low frequency) and HH (high frequency-high frequency) four subbands. For each subband, the energy value (i.e. the sum of the squares of all coefficient absolute values) is calculated, and the energy of the four subbands is normalized to generate a four-dimensional energy vector.

[0042] On this basis, two indexes are defined: one is the energy concentration degree, that is, the proportion of the maximum energy subband; the second is the energy ratio feature, that is, the energy ratio between each high frequency subband, which is used to identify the local texture directionality and frequency distribution form. This path can be used to accurately describe the detail distribution and direction response law of the image block under different frequency bands.

[0043] Finally, the above three types of indexes: frequency domain sparsity, spatial gradient density, and wavelet energy distribution are integrated into a compressed texture feature vector. Each image block corresponds to an independent feature vector, which is used as the core input basis for subsequent construction of compressed texture index atlas and image block reconstruction priority ranking.

[0044] The three-path joint extraction mechanism can accurately describe the image block attributes from the frequency response, spatial structure and multi-scale texture, which makes up for the deficiency of traditional image division and coarse-grained description method that cannot capture the local complexity difference in detail, and further improves the controllability, stability and reconstruction efficiency of the AIGC generation path.

[0045] In some embodiments of the present application, when constructing the compressed texture index atlas, the following steps are included: A three-dimensional feature mapping space is established, and the frequency domain sparsity, spatial gradient density and wavelet energy distribution of the image block are mapped to three-dimensional coordinate axes respectively to form a compressed texture feature scatter plot; In the compressed texture feature scatter plot, a local density clustering algorithm is used to identify high-density feature clusters. By setting the mean value of feature density and the radius range of the clustering center neighborhood, the image blocks located in the high-density clustering and the three-dimensional coordinate difference within the preset tolerance threshold are marked to construct a compressed texture redundant block set; For the image blocks that are not marked as compressed texture redundant blocks, the Euclidean distance between the coordinate points corresponding to the frequency domain sparsity, spatial gradient density and wavelet energy distribution in the three-dimensional feature mapping space and the feature distribution boundary is calculated, and the reconstruction priority score is generated according to the Euclidean distance and the spatial position boundary of the image block in the original image source. Sort all image blocks according to the reconstruction priority score from high to low, construct a compressed texture index atlas, and the compressed texture index atlas includes the spatial coordinates, frequency domain sparsity, spatial gradient density, wavelet energy distribution and corresponding reconstruction priority score of the image blocks, which are used as the basis for guiding the AIGC generation path.

[0046] In the embodiment, in order to establish an effective guiding mechanism for the image block generation order, improve the structural integrity and detail retention capability of the AIGC image generation, a compressed texture index atlas construction method based on the combination of multi-dimensional feature clustering and spatial position evaluation is proposed. The core of the method is to construct a three-dimensional feature mapping space through three types of compressed texture indicators (i.e. frequency domain sparsity, spatial gradient density, and wavelet energy distribution), and to identify redundant areas and determine the reconstruction priority in the space, and finally realize the ordered planning of the image block generation path. The specific steps include: (1) Constructing a three-dimensional feature mapping space and a feature scatter plot: first, take the three indicators in the compressed texture feature vector of each image block as the three coordinate axes of the three-dimensional coordinate system: the X-axis corresponds to the frequency domain sparsity, the Y-axis corresponds to the spatial gradient density, and the Z-axis corresponds to the wavelet energy distribution. Draw feature points in the three-dimensional coordinate system in units of image blocks to form a compressed texture feature scatter plot. The plot is used to visually represent the differences and clustered distribution of image blocks in texture attributes.

[0047] (2) Identifying a set of compressed texture redundant blocks: perform a density-based clustering operation on the above feature scatter plot. Preferably, use local density clustering (such as LOF or DBSCAN), which can identify high-density clusters based on the neighborhood density of points. Specifically: first, calculate the average local density of all points in the entire feature space as a global density reference; then set a neighborhood radius range at each high-density cluster center, and count the density and coordinate difference of the points in the neighborhood; if a three-dimensional coordinate point of an image block is located in the neighborhood of a cluster center with a density higher than the average value, and the distance between the cluster center in the three dimensions is within the preset tolerance threshold, then mark the image block as a compressed texture redundant block. Such image blocks often have highly repeated texture features and low reconstruction information contribution, and can be delayed or compressed through image fusion to save computing resources.

[0048] It should be noted that the neighborhood radius range can be obtained by the following method: (1) Calculate the Euclidean distance between all image block points in the three-dimensional mapping space, and construct a complete distance matrix. (2) Calculate the average distance and standard deviation of all point pairs, extract all non-repeated point pair distances from the distance matrix, and calculate the global average distance μ and global standard deviation σ (3) Construct the neighborhood radius range R with μ and σ Preferably, the neighborhood radius range R is set as: R = μ ± β × σ wherein the global average distance μ is the average mutual distance of the overall feature points; the global standard deviation σ is the diffusion measure of the overall distribution; β is the regulation coefficient, the value range is set to [0.8, 1.5], which can be adjusted according to the training set or experimental results.

[0049] (3) Evaluate the reconstruction priority score of the remaining image blocks: For image blocks that are not marked as compressed texture redundant blocks, the urgency and priority of their reconstruction in the image generation process need to be evaluated. Specifically, it includes two indicators: Feature distance indicator: In the three-dimensional feature space, the Euclidean distance between each image block corresponding point and the feature scatter boundary is calculated. This distance reflects whether the image block has significant texture difference, the greater the distance, the more unique and reconstructive the texture features of the image block.

[0050] Spatial location indicator: Combined with the two-dimensional spatial coordinates of the image block in the original image source, it is judged whether it is located in the edge area of the image, the center of the main structure or the vicinity of the important semantic area. If the image block is located in the key area such as structure fracture zone and color jump zone, its reconstruction priority score is increased.

[0051] The above two scoring factors are combined according to the weighting strategy to generate the image block reconstruction priority score for sorting control.

[0052] (4) Build a compressed texture index map: all image blocks that are not marked as redundant are arranged in order of high to low according to the reconstruction priority score to form a compressed texture index map. The map is a structured data table that records the spatial coordinates, frequency domain sparsity, spatial gradient density, wavelet energy distribution and reconstruction priority score of each image block. The system can process them one by one in the order of the map when performing AIGC image block generation, ensuring that the key structure area is generated first and the texture redundant area is processed later.

[0053] Through the index map, the dynamic scheduling of image blocks in the AIGC generation path can be realized, the structure recovery efficiency and the overall coherence of the generated image are improved, and an intelligent image reconstruction scheduling mechanism with structure difference perception and texture redundancy control is constructed.

[0054] In some embodiments of the present application, when generating a semantic residual mapping map, it includes: Based on the pyramid level to which the image block belongs, a local image area with an area of one tenth of the original image area is extracted in the original image source as the center of the image block, as the basic input of the regional semantic background map; Input the image block and the regional semantic background map into the same semantic analysis model respectively, perform semantic segmentation and semantic boundary extraction operation, extract semantic label matrix, boundary contour map and structure level annotation map; The semantic label matrix of the image block and the region semantic background graph is compared in difference, to generate a semantic coverage difference graph; The boundary contour graph of the image block and the region semantic background graph is executed structure contour comparison, and a boundary offset graph is generated based on boundary direction similarity, intersection length and closure degree index; The hierarchical matching of the structure level label graph of the image block and the region semantic background graph is executed, the hierarchical difference and the nesting error are counted, and a hierarchical deviation graph is generated. The semantic coverage difference graph, the boundary offset graph and the hierarchical deviation graph are registered and superimposed in image coordinate position, to generate a semantic residual mapping graph.

[0055] In this embodiment, in some embodiments of the present application, the processing steps of generating a semantic residual mapping graph, based on the pyramid level information of the image block, are embedded into the semantic context of a larger region for difference comparison, which includes the following contents: Firstly, based on the pyramid level to which the image block belongs, a local region centered on the image block is intercepted in the original image source, and the area of the local region is set to one tenth of the area of the original image to ensure that it contains sufficient context information as the basis for input of the region semantic background graph. The area selection ratio is set according to the experiment, aiming to balance the integrity of the context information and the calculation efficiency.

[0056] Subsequently, the image block and the corresponding region semantic background graph are input into a unified semantic analysis model at the same time. The semantic analysis model can be a lightweight semantic segmentation network (such as DeepLabv3+ or SegFormer), which is used to extract three types of structured semantic information, including: semantic label matrix: representing the semantic category corresponding to each pixel in the image; boundary contour graph: based on semantic boundary extraction operation, identifying the edge direction between adjacent semantic regions; structure level label graph: according to the nesting and relative position relationship of the semantic region, labeling the structure level information in the image, such as background-foreground-object components, etc.

[0057] After obtaining the above three types of structures, three residual comparison analyses are carried out in turn: Semantic label matrix difference analysis: compare the semantic label matrix of the image block and the corresponding region of the background graph at the pixel level, calculate the inconsistent area of the semantic category, and label the semantic coverage difference graph. This graph records the spatial position and category number difference of the semantic missing or misplacement area.

[0058] Boundary contour graph structure comparison analysis: the boundary contour graphs in the image block and the background graph are compared in boundary structure, the boundary direction similarity (which can be calculated by the main direction consistency rate), the intersection length error (which is measured by the boundary overlap ratio) and the boundary closure degree (which judges whether the boundary forms a closed region) are quantified, and finally a boundary offset graph is generated.

[0059] Structure hierarchy annotation map comparison analysis: based on the structure hierarchy annotation map, the hierarchical nesting structure of the same semantic labels in the two regions is paired, the hierarchical depth difference and the nesting structure damage degree are counted, and a hierarchy deviation map is output.

[0060] Finally, the above three maps (semantic coverage difference map, boundary offset map and hierarchy deviation map) are pixel-level registered and superimposed according to the original image coordinate position, and a unified semantic residual mapping map is generated. The map is an important reference for the subsequent semantic completion and structure repair generation module, and has structure guiding and content difference representation ability.

[0061] In some embodiments of the present application, the step of performing semantic structure analysis on each image block region comprises: performing semantic segmentation on the image block using an image segmentation network to obtain a corresponding semantic label matrix; extracting a boundary contour map according to the semantic label matrix, and aligning the edges in combination with the grayscale map of the image block; annotating the bounding rectangle of each semantic region in the semantic label matrix, and counting the boundary contact length and semantic class combination mode between adjacent semantic regions; constructing a structure hierarchy annotation map based on the boundary contact relationship, the structure hierarchy annotation map being used to represent the hierarchical nesting relationship and local structure topology of the semantic regions in the image block.

[0062] In this embodiment, first, the target image block is subjected to semantic segmentation processing using an image segmentation network. The image segmentation network can be SegFormer, DeepLabv3+, HRNet or other network structure with multi-scale perception ability and semantic boundary preservation ability. The above segmentation network can assign a clear semantic class label to each pixel in the image block, forming a semantic label matrix consistent with the size of the input image block. Each position value in the semantic label matrix represents the semantic class number of the corresponding pixel, which is used for subsequent structure information analysis.

[0063] Next, the boundary contour information between the semantic regions is extracted from the generated semantic label matrix to form a boundary contour map. This process includes: based on the jump region of the semantic class in the label matrix, the class boundary is detected using the Laplacian operator or the Canny operator; then, in combination with the grayscale map corresponding to the image block, the pixel gradient of the grayscale gradient and the semantic jump region is compared, and the sub-pixel level alignment optimization of the edge position is performed, to ensure that the semantic boundary and the actual physical edge of the image coincide as much as possible. This operation can improve the accuracy of structure recognition and reduce the error caused by semantic misplacement.

[0064] Subsequently, in the above semantic label matrix, the minimum circumscribed rectangle frame of each independent semantic region is extracted. By counting the boundary contact length between any two adjacent semantic regions, the boundary contact mode (such as straight line type, arc or angle), and the semantic category combination mode (such as “object-background”, “foreground-foreground” and the like), a structural contact relationship diagram between the semantic regions is established.

[0065] On this basis, a structural hierarchical annotation diagram is constructed to represent the nesting relationship and local topological structure of each semantic region in the image block. Specifically, by analyzing the inclusion relationship of the circumscribed rectangle and the relative position of the center point, the surrounding and surrounded relationship between the semantic regions is classified to determine whether it constitutes a nested structure. For example: when the minimum circumscribed rectangle of a certain semantic region completely contains another region, and the center point falls inside, it can be determined that it is a hierarchical nesting relationship. At the same time, combined with the boundary contact relationship diagram, a hierarchical representation framework with topological structure can be established. The structural hierarchical annotation diagram encodes the inclusion, adjacency and common edge relationship between the semantic regions in the form of a graph structure, providing a structural reference basis for the subsequent semantic residual calculation and generation model.

[0066] In some embodiments of the present application, the step of performing semantic residual calculation on the original image block and the regional semantic background image includes: The original image block and the regional semantic background image are respectively encoded into vector representation, and the semantic category feature distribution is extracted using a word embedding module; Based on the semantic nesting level recorded in the structural hierarchical annotation diagram, the two vector representations are segmented and paired; Cosine similarity calculation is performed on each paired segment to construct a multi-segment semantic matching matrix; In the multi-segment semantic matching matrix, the unaligned segment sequence is extracted, and through structural consistency analysis and color texture mutation recognition, it is labeled as a semantic residual segment; All semantic residual segments are reconstructed into a semantic residual mapping diagram, which includes semantic segment position, semantic category, structural disconnection factor and residual confidence score.

[0067] In this embodiment, the original image block and its corresponding regional semantic background image are respectively input into a unified semantic encoding module for feature encoding to form a semantic vector representation. The semantic encoding module can use a pre-trained semantic embedding network, such as a word vector model (such as Word2Vec, GloVe or BERT embedding) combined with the semantic label output by the image semantic segmentation network, to map each semantic label to a vector representation, and then combine the proportion of each semantic label in the region to obtain a weighted average of the entire image to obtain an image-level semantic category feature distribution vector. This processing realizes the unified representation of the semantic content of the image block in the embedding space.

[0068] The semantic vector segments are segmented and paired according to the nested level of the semantic regions in the background graph, that is, the image blocks and the semantic regions at the same level in the background graph are extracted respectively to form corresponding semantic vector segments. Each paired segment represents a set of structurally comparable semantic information units.

[0069] For each set of semantic vector segments, a cosine similarity calculation is performed to measure the semantic consistency of the original image block and the region semantic background graph at the same semantic level, obtaining a multi-segment semantic matching matrix composed of a set of similarity scores. Each row of the matrix corresponds to the semantic level of the original image block, and the column corresponds to the region semantic background graph. Each element in the matrix is the matching score of the corresponding semantic segment.

[0070] In the semantic matching matrix, all segment pairs below the set threshold are screened out and identified as unaligned segment sequences. These segments may have abnormal expression due to local structural mutations, class mismatches or missing. To accurately define the semantic residual, structural consistency analysis of these unaligned segments is also required, that is, the spatial topology, nesting relationship, boundary direction and other attributes recorded in the structural hierarchy annotation graph are compared to determine the degree of structural disconnection. At the same time, the color and texture mutation recognition mechanism of the image block and the background graph in this segment area is introduced to detect whether it reaches the abnormal threshold in terms of color histogram difference and texture spectrum difference.

[0071] For segments with severe structural consistency deviation or prominent color and texture changes, they are uniformly labeled as semantic residual segments. All semantic residual segments are reconstructed, and based on their spatial coordinates, semantic categories, structural consistency evaluation results and color mutation scores, a semantic residual mapping graph containing four types of information is generated, that is: Semantic segment position, used to identify the specific spatial coordinate range of each semantic residual segment in the image block, the acquisition process includes: Locate the semantic segments that fail to complete pairing in the semantic matching matrix; backtrack to the semantic label matrix in the original image block and the region semantic background graph to extract the mask area of the segment in the two-dimensional image space; calculate the bounding box of the mask area to obtain the top-left and bottom-right coordinates of the semantic residual segment in the image, or record its spatial distribution in the form of a mask for subsequent spatial encoding in the input tensor.

[0072] Semantic category, used to indicate the category of the residual fragment in the semantic parsing model, the acquisition method is as follows: from the semantic label matrix of the original image block and the regional semantic background image respectively, the residual mask region is extracted; if the semantic labels are different, the label of the region in the original image block is used as the semantic category; all semantic categories are encoded into fixed dimension vectors (such as one-hot or integer encoding) through a preset dictionary, and written into the prompt information input tensor.

[0073] Structural dislocation factor, used to quantify the deviation of the fragment at the structure level, the acquisition method is as follows: in the structure level annotation map, the nested path of the semantic unit to which the residual fragment belongs in the original image and the background image is extracted (i.e. its level and parent node chain); if the nested path is different, calculate the path length difference, the least common ancestor offset and the level dislocation number; the above three indicators are standardized and weighted to obtain a floating point score as the structural dislocation factor, which is used to represent the abnormality of the semantic fragment in the structure topology.

[0074] Residual confidence score, reflecting the influence of the fragment on the overall semantic distortion, the acquisition method includes: for each unpaired semantic fragment, calculate the cosine similarity and cross entropy loss between it and the corresponding background region in the semantic matching matrix; on the texture level, perform gradient statistics and texture direction divergence calculation on the region boundary to obtain the texture mutation degree; normalize and fuse the semantic difference and the texture mutation degree, and generate the residual confidence score according to the weighted result; high-score fragments are preferentially used for attention weighted input in the generation stage to improve the pertinence of key area repair.

[0075] In some embodiments of the present application, the step of inputting the semantic residual mapping into the preset AIGC generation model to complete the image block content generation includes: Based on the semantic residual mapping, a prompt information input tensor is constructed, which contains the spatial coordinates, semantic category code, residual confidence score and structural dislocation factor of each semantic residual fragment; According to the pyramid level to which the image block belongs, the corresponding generation sub-model is called, and the original image block and the prompt information input tensor are input simultaneously in a multi-channel input mode; An attention routing mechanism is introduced in the AIGC generation model, and the activation weight of the multi-layer attention mechanism is adjusted according to the semantic category code and the residual confidence score in the prompt information input tensor; After generating the output image block, the output image block and the original image block are compared for local coincidence degree, to verify whether the semantic recovery degree and the structural continuity meet the standard, and if not, the semantic residual mapping is re-input for generation update.

[0076] In this embodiment, first, based on the information recorded in the aforementioned semantic residual map, a prompt information input tensor is constructed. This tensor serves as a model-assisted input signal to guide the generative model to perform semantic completion and structural repair in key areas. Specifically, the prompt information input tensor is initialized with the same spatial dimensions as the image block and is encoded layer by layer in the following channels: the first channel, the spatial coordinate information of the semantic residual segment, which uses position encoding to mark the residual area; the second channel, semantic category encoding, which uses a fixed dictionary to assign independent encoding to each semantic class; the third channel, residual confidence score, which expresses the semantic residual severity of each area in the form of a floating-point number; the fourth channel, structural disconnection factor, which combines the nested deviation quantification results obtained from structural consistency analysis.

[0077] The construction method of the prompt information tensor not only preserves the spatial distribution information of semantics and structure, but also embeds the semantic importance difference through numerical encoding, guiding the generative network to focus on the key fragments of semantic residual.

[0078] Subsequently, according to the pyramid level to which the image block belongs, the corresponding resolution level of the generative sub-model is called from the pre-set multi-level generative sub-model library. Each generative sub-model in this model library can be a lightweight variant based on U-Net, Diffusion or Transformer architecture, optimized for the texture details and semantic complexity of different pyramid levels.

[0079] During input, a multi-channel input structure is used to input the original image block and the prompt information tensor into the AIGC generative model simultaneously. In the encoding stage, the semantic information in the prompt tensor is mapped to the image feature space through cross-attention mechanism, realizing the dynamic perception of the key areas by the generative path.

[0080] To further improve the response degree of the model to the prompt signal, an attention routing mechanism is introduced into the model: according to the semantic category encoding and residual confidence score in the prompt tensor, the activation weights of different channels in each layer of attention module are dynamically adjusted. For example, for areas with high residual confidence, the weights of the channels corresponding to their semantic categories are increased to encourage the model to strengthen the expression and reconstruction of these fragments during generation.

[0081] After generating the output image block, a local coincidence comparison is performed with the original image block to evaluate whether the semantic recovery degree and structural continuity meet the preset standard. The comparison methods include: based on the structure level annotation map, checking whether the generated image retains the original nesting relationship; performing cosine similarity and cross entropy comparison on the semantic label reconstruction region to confirm semantic consistency; performing texture change analysis on the boundary fusion area to ensure that there is no abrupt jump or edge misplacement phenomenon. If the detection result does not meet the threshold condition, the original semantic residual error map is automatically re-input into the model to trigger the generation update mechanism. The mechanism can introduce adjusted attention path, enhance residual error weight or call redundant alternative model strategy to ensure that the semantic and structural quality of the generated image block meets the standard, and finally used for image fusion step.

[0082] In some embodiments of the present application, the step of image quality evaluation includes: The preliminary synthesized image is analyzed in correspondence with the original image source at the image block level to calculate the texture continuity score, color consistency score and structural integrity score; The texture continuity score is obtained by counting the normalized value of the frequency response difference in the edge region of adjacent image blocks; The color consistency score is obtained by comparing the Bhattacharyya distance of the RGB color channel histogram in the boundary neighborhood of each image block; The structural integrity score is evaluated by counting whether the spatial connectivity and nested structure of the region semantic label matrix in the reconstructed image are consistent; The three types of scores form an image generation stability index group, which is used as a judgment basis for iterative optimization of the final image generation result.

[0083] In this embodiment, each residual segment in the semantic residual error map is attached with the following four types of information, namely semantic segment position, semantic category, structure dislocation factor and residual error confidence score, and the specific acquisition process is as follows: The semantic segment position is used to identify the specific spatial coordinate range of each semantic residual segment in the image block, and the acquisition process includes: locating the semantic segment that fails to complete pairing in the semantic matching matrix; backtracking to the semantic label matrix in the original image block and the region semantic background map to extract the mask area of the segment in the two-dimensional image space; calculating the bounding box of the mask area to obtain the top-left corner and bottom-right corner coordinates of the semantic residual segment in the image, or recording its spatial distribution in the form of mask for spatial coding in the subsequent input tensor.

[0084] Semantic category, used to indicate the category belonging of the residual fragment in the semantic parsing model, the acquisition method is as follows: from the semantic label matrix of the original image block and the regional semantic background image respectively, the residual mask region is extracted; if the semantic labels are different, the label of the region in the original image block is used as the semantic category; all semantic categories are encoded into fixed dimension vectors (such as one-hot or integer encoding) through a preset dictionary, and written into the prompt information input tensor.

[0085] Structural dislocation factor, used to quantify the deviation degree of the fragment at the structure level, the acquisition method is as follows: in the structure level annotation map, the nested path (i.e. the level and parent node chain) of the semantic unit to which the residual fragment belongs in the original image and the background image is extracted; if the nested path is different, the path length difference, the least common ancestor offset and the level dislocation number are calculated; the above three indexes are standardized and weighted to obtain a floating point score as the structural dislocation factor, which is used to represent the abnormal degree of the semantic fragment in the structure topology.

[0086] Residual confidence score, reflecting the influence degree of the fragment on the overall semantic distortion, the acquisition method includes: for each unpaired semantic fragment, the cosine similarity and cross entropy loss between it and the corresponding background region in the semantic matching matrix are calculated; on the texture level, gradient statistics and texture direction divergence calculation are performed on the boundary of the region to obtain the texture mutation degree; the semantic difference degree and the texture mutation degree are normalized and fused to generate the residual confidence score according to the weighted result; high score fragments are preferentially used for attention weighted input in the generation stage to improve the pertinence of key area repair.

[0087] Referring to Figure 2 As shown in the figure, the embodiment of the application provides an AIGC image fast generation system, which comprises: An image block construction module is configured to obtain an original image source, divide the original image source into image blocks, construct an image block set based on a multi-layer pyramid rule, and extract multi-scale texture compression features of each image block, including frequency domain sparsity, spatial gradient density and wavelet energy distribution. A texture index map construction module is configured to construct a compressed texture index map based on the multi-scale texture compression features, sort the compressed texture index map according to the reconstruction priority of the image block, and use the compressed texture index map as a guide basis for the AIGC generation path. A semantic residual mapping module is configured to perform semantic structure analysis on each image block region, construct a corresponding regional semantic background image, and perform semantic residual calculation on the original image block and the regional semantic background image to generate a semantic residual mapping image. An image generation module is configured to input the semantic residual mapping image into a preset AIGC generation model to complete the generation of the content of the target image block. The boundary fusion module is configured to perform boundary fusion processing on the plurality of generated image blocks, adopts an edge jump detection and color histogram matching strategy, performs dynamic fusion on edge regions, and generates a preliminary synthesized image; The quality evaluation module is configured to perform image quality evaluation on the preliminary synthesized image, and extract image generation stability indexes including texture continuity, color consistency, and structural integrity. The feedback optimization module is configured to feed back the image generation stability index results to the construction process of the compressed texture index atlas and the semantic residual mapping graph, update the reconstruction priority ranking and residual information extraction strategy, perform iterative optimization processing on the image reconstruction path, and output the final image generation result.

[0088] Finally, it should be noted that: the above examples are only used to illustrate the technical solutions of the present application, but not to limit it, although the present application has been described in detail with reference to the above examples, those skilled in the art should understand that: the specific embodiments of the present application can be modified or replaced by the same, without departing from the spirit and scope of the present application, any modification or equivalent replacement, which should be covered within the protection scope of the claims of the present application.

Claims

1. A method for rapid generation of AIGC images, characterized in that, include: The original image source is obtained, and the original image source is divided into image patches. An image patch set is constructed based on the multi-level pyramid rule, and the multi-scale texture compression features of each image patch are extracted, including frequency domain sparsity, spatial gradient density and wavelet energy distribution. A compressed texture index map is constructed based on multi-scale texture compression features, and sorted according to the reconstruction priority of image patches. The compressed texture index map is used as the guiding basis for the AIGC generation path. Semantic structure parsing is performed on each image block region to construct the corresponding region semantic background map, and semantic residual calculation is performed on the original image block and the region semantic background map to generate a semantic residual mapping map; The semantic residual map is input into the preset AIGC generation model to generate the content of the target image patch; Multiple generated image patches are subjected to boundary fusion processing. An edge transition detection and color histogram matching strategy are used to dynamically fuse edge regions and generate a preliminary synthetic image. Image quality assessment is performed on the preliminary synthesized image, and image generation stability indicators including texture coherence, color consistency and structural integrity are extracted. The image generation stability index results are fed back into the construction process of the compressed texture index map and semantic residual map, the reconstruction priority ranking and residual information extraction strategy are updated, the iterative optimization process of the image reconstruction path is performed, and the final image generation result is output.

2. The AIGC image fast generation method according to claim 1, characterized in that, When dividing the original image source into image patches and constructing an image patch set based on multi-level pyramid rules, the following is included: After obtaining the original image source, the original image is initially divided according to a fixed size to obtain multiple image blocks; Perform the following operations for each image block: The high-frequency subband of the image patch is extracted using discrete wavelet transform, and the energy density value in the high-frequency subband is calculated to characterize the texture complexity. Perform Sobel edge detection on the image patch, count the gradient direction of all edge pixels, calculate the ratio of the number of occurrences of the main direction to the number of occurrences of other directions, and obtain the edge direction consistency score. Calculate the mean and standard deviation of the three RGB color channels of the image block respectively, and calculate the sum of the variances between the color channels to characterize the degree of difference in color composition; Texture complexity, edge direction consistency score, and color composition difference are used as feature vectors and input into a preset multi-level pyramid mapping rule. The mapping rule matches the feature vectors with the reference threshold range of each pyramid level in the multi-dimensional feature space to determine the pyramid level to which the image patch belongs. Image patches belonging to the same pyramid level in the judgment results are grouped together, and finally a multi-level pyramid structure containing multiple levels of image patch sets is constructed.

3. The AIGC image fast generation method according to claim 2, characterized in that, When extracting multi-scale texture compression features from each image patch, the following steps are included: For each image patch, three feature extraction paths are used: (1) Frequency domain sparsity extraction path: Perform discrete Fourier transform on the image patch to obtain the frequency domain coefficient distribution, and count the proportion of frequency domain coefficients with amplitude greater than the preset amplitude, which is recorded as frequency domain sparsity; (2) Spatial gradient density extraction path: The Sobel operator is used to extract the gradient image of the image block in the horizontal and vertical directions. The pixels of the gradient image are counted, the proportion of pixels whose gradient magnitude exceeds the set gradient threshold is counted, and the gradient direction variance is calculated to form the spatial gradient density. (3) Wavelet energy distribution extraction path: Based on the Daubechies wavelet, perform three-level wavelet decomposition on the image block to obtain LL, LH, HL and HH subbands, normalize the energy of each subband, construct the energy distribution vector, and use the energy concentration and subband energy ratio to represent the wavelet energy distribution characteristics; Frequency domain sparsity, spatial gradient density, and wavelet energy distribution are combined into a compressed texture feature vector, which serves as the basic input for the construction and priority ranking of the compressed texture index map.

4. The AIGC image fast generation method according to claim 3, characterized in that, When constructing a compressed texture index map, the following steps are included: A three-dimensional feature mapping space is established, and the frequency domain sparsity, spatial gradient density and wavelet energy distribution of image patches are mapped to the three-dimensional coordinate axes respectively to form a compressed texture feature scatter plot. In the compressed texture feature scatter plot, a local density clustering algorithm is used to identify high-density feature clusters. By setting the mean feature density and the radius range of the neighborhood of the cluster center, image blocks located inside the high-density clusters and whose three-dimensional coordinate differences are within a preset tolerance threshold are marked to construct a set of compressed texture redundant blocks. For image patches that are not marked as compressed texture redundancy blocks, calculate the frequency domain sparsity, spatial gradient density and wavelet energy distribution corresponding coordinate points and the feature distribution boundary in the three-dimensional feature mapping space. Based on the Euclidean distance and the spatial location boundary of the image patch in the original image source, generate a reconstruction priority score. All image patches are sorted from high to low according to their reconstruction priority scores, and a compressed texture index map is constructed. The compressed texture index map includes the spatial coordinates, frequency domain sparsity, spatial gradient density, wavelet energy distribution and corresponding reconstruction priority scores of the image patches, which are used as the guiding basis for the AIGC generation path.

5. The AIGC image fast generation method according to claim 4, characterized in that, When generating the semantic residual map, the following steps are included: Based on the pyramid level to which the image patch belongs, a local image region centered on the corresponding position of the image patch and covering an area of ​​one-tenth of the original image area is extracted from the original image source and used as the basic input for the region semantic background map. Image patches and region semantic background maps are respectively input into the same semantic parsing model to perform semantic segmentation and semantic boundary extraction operations, and extract semantic label matrix, boundary contour map and structural hierarchy annotation map. The semantic label matrices of image patches and region semantic background maps are compared to generate a semantic coverage difference map. Perform structural contour comparison on the boundary contour maps of image patches and region semantic background maps, and generate boundary offset maps based on boundary orientation similarity, intersection length and closure index; Perform hierarchical matching on the structural hierarchy annotation maps of image patches and region semantic background maps, calculate the hierarchical difference and nesting error, and generate a hierarchical deviation map. The semantic coverage difference map, boundary offset map, and hierarchical deviation map are registered and overlaid at image coordinate positions to generate a semantic residual mapping map.

6. The AIGC image fast generation method according to claim 5, characterized in that, The steps for semantic structure parsing of each image patch region include: An image segmentation network is used to perform semantic segmentation on image patches to obtain the corresponding semantic label matrix; Boundary contour maps are extracted based on the semantic tag matrix, and edge alignment is performed by combining the grayscale images of image patches; Mark the bounding rectangle of each semantic region in the semantic label matrix, and calculate the boundary contact length between adjacent semantic regions and the semantic category combination pattern; A structural hierarchy annotation map is constructed based on boundary contact relationships. The structural hierarchy annotation map is used to represent the hierarchical nesting relationship and local structural topology of each semantic region in an image block.

7. The AIGC image fast generation method according to claim 6, characterized in that, The steps for calculating the semantic residual between the original image patch and the region semantic background map include: The original image patch and the region semantic background map are encoded into vector representations respectively, and the semantic category feature distribution is extracted using a word embedding module; Based on the semantic nesting hierarchy recorded in the structural hierarchy annotation diagram, the vector representations of the two are segmented and paired. Perform cosine similarity calculation on each paired segment to construct a multi-segment semantic matching matrix; Unaligned segment sequences are extracted from multiple semantic matching matrices and labeled as semantic residual segments through structural consistency analysis and color texture mutation identification. All semantic residual fragments are reconstructed into a semantic residual map, which includes the semantic fragment location, semantic category, structural disjointness factor, and residual reliability score.

8. The AIGC image fast generation method according to claim 7, characterized in that, The steps for generating image patch content by inputting the semantic residual map into the preset AIGC generation model include: A prompt information input tensor is constructed based on the semantic residual mapping graph. The tensor contains the spatial coordinates, semantic category code, residual reliability score and structural disjoint factor of each semantic residual segment. Based on the pyramid level to which the image patch belongs, the corresponding generation sub-model is called, and the original image patch and the prompt information input tensor are input simultaneously using a multi-channel input method. An attention routing mechanism is introduced into the AIGC generative model, which adjusts the activation weights of the multi-layer attention mechanism based on the semantic category encoding and residual reliability score in the input tensor of the prompt information. After generating the output image patch, the local overlap between the output image patch and the original image patch is compared to verify whether the semantic recovery and structural coherence meet the standards. If they do not meet the standards, the semantic residual map is re-inputted for generation and updating.

9. The AIGC image fast generation method according to claim 8, characterized in that, The steps involved in image quality assessment include: The preliminary synthesized image is compared with the original image source at the image patch level to calculate the texture coherence score, color consistency score and structural integrity score. Texture coherence score is obtained by statistically analyzing the normalized values ​​of frequency response differences in the edge regions of adjacent image patches; Color consistency score is obtained by comparing the Bhattacharyya distance of the RGB color channel histograms of each image patch in the boundary neighborhood; Structural integrity score is evaluated by assessing whether the spatial connectivity of the statistical region semantic label matrix in the reconstructed image is consistent with the nested structure; The three types of scores form a set of image generation stability indicators, which serve as the basis for judging the iterative optimization of the final image generation result.

10. A rapid AIGC image generation system, used to implement the method according to any one of claims 1-9, characterized in that, The system includes: The image patch construction module is configured to acquire the original image source, divide the original image source into image patches, construct an image patch set based on the multi-level pyramid rule, and extract the multi-scale texture compression features of each image patch, including frequency domain sparsity, spatial gradient density and wavelet energy distribution. The texture index map construction module is configured to construct a compressed texture index map based on multi-scale texture compression features, sort the image patches according to their reconstruction priority, and use the compressed texture index map as the guiding basis for the AIGC generation path. The semantic residual mapping module is configured to perform semantic structure parsing on each image block region, construct the corresponding region semantic background map, and perform semantic residual calculation on the original image block and the region semantic background map to generate a semantic residual mapping map. The image generation module is configured to input the semantic residual map into a preset AIGC generation model to generate the content of the target image patch; The boundary fusion module is configured to perform boundary fusion processing on multiple generated image patches. It adopts edge transition detection and color histogram matching strategies to dynamically fuse edge regions and generate a preliminary synthesized image. The quality assessment module is configured to perform image quality assessment on the initial synthesized image and extract image generation stability indicators including texture coherence, color consistency and structural integrity. The feedback optimization module is configured to feed back the image generation stability index results to the construction process of the compressed texture index map and semantic residual map, update the reconstruction priority ranking and residual information extraction strategy, perform iterative optimization processing of the image reconstruction path, and output the final image generation result.

Citation Information

Cited By

  • Method for finely adjusting color difference of adjacent pictures in text H5 scene

    CN121582359A