Multi-modal data synthesis system based on generative adversarial network

Through the cooperation of dynamic hierarchical division and interactive selection units, combined with multimodal feature alignment and stage seamless splicing technology, the problems of regional positioning blur and multimodal fusion deviation in the existing system are solved, and high-precision area adjustment and high-quality image synthesis are achieved.

CN120070207AActive Publication Date: 2025-05-30CHENGDU POLYTECHNIC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510542303.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The existing multimodal data synthesis system based on generative adversarial networks has problems such as regional positioning, limitations of global optimization and multimodal fusion bias, resulting in the offset of the synthesis content and the generation results that are inconsistent with the reference graph style.

Method used

The dynamic hierarchical division unit is used to divide the reference map recursively based on image feature points, combined with the interactive selection unit, allows users to mark the synthetic areas to be adjusted and enter text descriptions, and jointly encode text and image features through the multimodal feature alignment unit, and finally local generation and multimodal fusion are achieved through the stage seamless splicing unit.

Benefits of technology

Improve the accuracy of regional positioning, reduce the offset of synthetic content, and achieve finer local adjustments and higher quality synthetic images, ensuring that the generated results are consistent with the reference picture style.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070207A_ABST
    Figure CN120070207A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to a multi-modal data synthesis system based on a generative adversarial network. The system comprises a dynamic hierarchical division unit, an interactive selection unit, a multi-modal feature alignment unit and a stage seamless splicing unit. According to the method, the limitation of global optimization is avoided, finer local adjustment is allowed, a user can mark a specific area and input description, personalized requirements are met, the user can select time or precision priority according to requirements on the premise that the consistency of text and image features is ensured and the quality and style consistency of a generated result are enhanced, and the user experience is improved. And the method adapts to different use scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and more specifically, to a multi-modal data synthesis system based on a generative adversarial network. Background Art

[0002] AI image synthesis technology refers to the technology of using artificial intelligence algorithms, especially machine learning and deep learning methods, to generate, modify or synthesize images, which enables a computer to understand and generate complex visual content, and realizes various applications from simple image editing tasks to complex scene creation. Common AI image synthesis technologies include, but are not limited to, generative adversarial networks, which consist of a generative network and a discriminative network. Through the game process between the two, the generative network can learn how to generate realistic or appropriate images. The existing technology has been applied to many fields, such as super-resolution image generation, image conversion, and face generation, etc. Currently, when users perform AI image synthesis, they usually send a reference image to an AI program and input text. The generative network converts the text description into an image and synthesizes it at the corresponding position of the reference image. After the discriminative network determines, the final image is automatically generated. However, the current multi-modal data synthesis system based on generative adversarial networks has the following problems: Fuzzy region localization: It is unable to accurately identify the image region corresponding to the user's text description, resulting in the deviation of the synthesized content. Limitations in global optimization: Traditional methods perform unified feature extraction on the entire image, making it difficult to finely adjust local regions. Multi-modal fusion deviation: The alignment of text and image features is insufficient, and the generated result is inconsistent with the style of the reference image. In view of this, we propose a multi-modal data synthesis system based on generative adversarial networks. Summary of the Invention

[0003] The purpose of the present invention is to provide a multi-modal data synthesis system based on a generative adversarial network to solve the problem that the accuracy of AI in recognizing text and pictures is limited in the above-mentioned background art. Especially when making a separate adjustment to a certain area of the image, due to the inability to accurately identify the position and content corresponding to the text, the synthesized image has a deviation.

[0004] To achieve the above purpose, the present invention provides a multi-modal data synthesis system based on a generative adversarial network, including a dynamic hierarchical division unit, an interactive selection unit, a multi-modal feature alignment unit, and a stage seamless splicing unit; The dynamic hierarchical division unit recursively divides the reference image based on image feature points to form multi-level composite regions, which improves the accuracy of region positioning, reduces the offset of composite content, enables subsequent staged generation and stitching, avoids global optimization limitations, improves the fineness of local adjustment, and allows the interactive selection unit to allow users to mark the composite regions to be adjusted and input text descriptions for transforming and adjusting the composite regions to meet personalized needs, where: The multi-modal feature alignment unit is used to jointly encode the text description and the marked multi-level composite regions, enhances the consistency of text and image features, makes the generated result consistent with the style of the reference image, enables the stage seamless stitching unit to form candidate images through the local generative adversarial network in stages, and then seamlessly stitches with the reference image through the multi-modal fusion algorithm, ensuring the quality and consistency of the composite image.

[0005] As a further improvement of this technical solution, the dynamic hierarchical division unit includes an image feature extraction module and a region division module; The image feature extraction module is used to extract the feature points of the reference image by using Canny edge detection and Harris corner detection. During the clustering process, these feature points can be used as the basis for region division, helping to divide meaningful image regions, improving the accuracy of region positioning and the quality of subsequent image generation; The region division module randomly selects K initial centroids using the K-means algorithm, assigns each feature point to the nearest centroid, calculates the centroid of each cluster, repeats the assignment and update until the centroids no longer change to form multiple initial regions, further performs feature point detection and division on each initial region according to the hierarchy to form sub-regions until the set hierarchy is met, realizing the continuous refinement of the reference image, enabling the staff to more accurately lock the region according to the position to be changed, avoiding inaccuracies caused by automatic text recognition of the changed region, and improving the accuracy.

[0006] As a further improvement of this technical solution, the interactive selection unit includes an interaction feedback module and a text input module; The interaction feedback module is used to receive the hierarchical regions output by the region division module and display them on the interactive interface. The hierarchical regions include the initial reference image, initial regions, and sub-regions. When the current hierarchical region is not the region required by the user, the user is allowed to set the hierarchy until a marking signal is triggered. The user can mark multiple hierarchical regions to form multiple marking signals; After receiving the marking signal, the text input module types in text descriptions for the hierarchical regions corresponding to multiple marking signals to guide how to adjust the marked hierarchical regions, enabling the user to independently input how to adjust for each hierarchical region, which is beneficial for more precise local adjustment.

[0007] As a further improvement of the technical solution, the interaction feedback module further includes a sketch adjustment module, which is used to provide a hand-drawing tool on the interactive interface, allowing the user to use the hand-drawing tool to outline a custom area that needs to be adjusted in the initial reference image, so that the image feature extraction module perceives the custom area as the actual reference image for feature point extraction, which helps to avoid the high intensity caused by the simultaneous loading or running of multiple initial areas, is more targeted for subsequent marking, and improves the operation efficiency.

[0008] As a further improvement of the technical solution, when the multi-modal feature alignment unit performs joint encoding, it includes the following steps: Text embedding: Perform word segmentation processing by capturing a large amount of semantic information of the text, and encode the text to generate a text embedding vector. Use a pre-trained language model to generate a mapping relationship between the semantic information and the text embedding vector, and identify the semantic information captured by the text description typed by the text input module and input it into the pre-trained language model to output the text embedding vector; Image feature embedding: Perceive the image features of multiple hierarchical regions through the image feature extraction module to form feature vectors for each hierarchical region, and map them to the same semantic space as the text embedding through a fully connected layer; Joint encoding: Perform multi-head self-attention processing on the text embedding vector and the image feature vector respectively to enhance their respective semantic representations, and through the cross-attention mechanism, model the correlation between the text and image features corresponding to multiple hierarchical regions respectively, fuse the attention outputs of the text and the image to generate a final joint encoding vector. Each hierarchical region corresponds to a joint encoding vector. Through the multi-head self-attention and cross-attention mechanisms, the model can simultaneously capture the local details and global semantics of the text and the image, achieve efficient cross-modal alignment, and can gradually refine the correlation between the text and image features to generate more accurate joint encoding. At the same time, the result of the joint encoding can guide the generative adversarial network to adjust specific regions to meet the personalized needs of users.

[0009] As a further improvement of the technical solution, the stage seamless splicing unit includes a candidate image formation module and a multi-modal fusion module; The candidate image formation module is used to arrange the hierarchical regions in the order from local to whole, and sequentially receive the joint encoding vectors output by the multi-modal feature alignment unit to form a joint encoding vector sequence. Generate high-quality candidate images in sequence according to the generator of the local generative adversarial network, and only generate images of specific hierarchical regions instead of the whole image, which helps to improve the flexibility and fineness of generation. The discriminator distinguishes between the generated image and the real image; The multimodal fusion module is used to execute the multimodal fusion algorithm, ensuring that the text embedding and image feature embedding corresponding to the candidate image are mapped to the same space to form a fusion feature, ensuring that the two can be effectively combined, aligning the feature points of the generated candidate image with the reference image, and applying smoothing processing in the splicing area to ensure natural transition.

[0010] As a further improvement of this technical solution, the multimodal fusion module further includes an accuracy feedback module. The accuracy feedback module is used to establish a priority interaction key between accuracy and time, allowing the user to select accuracy priority or time priority, and triggering the seamless splicing posture of the reference image, including the following postures: Posture 1: When receiving the time priority signal, map multiple candidate images to the same space and synchronously fuse them to form a complete composite image; Posture 2: When receiving the accuracy priority signal, perform staged smooth splicing on the candidate images arranged in the order from local to global, and after each splicing, make the multimodal feature alignment unit re-jointly encode until a composite image is formed.

[0011] As a further improvement of this technical solution, the synchronous fusion to form a complete composite image includes the following steps: After the smooth splicing of multiple candidate images A at the lowest level a is completed to form a candidate image a1 at level a - 1, continue to perform smooth splicing on multiple candidate images a1 at level a - 1 until a composite image is formed; The staged smooth splicing of the candidate images arranged in the order from local to global, and after each splicing, making the multimodal feature alignment unit re-jointly encode until a composite image is formed, includes the following steps: After the smooth splicing of multiple candidate images A at the lowest level a is completed to form a spliced image A1 at level a - 1, make the multimodal feature alignment unit re-jointly encode the spliced image A1 with the text description to form a new candidate image a1, continue to perform smooth splicing on multiple candidate images a1 at level a - 1, and again make the multimodal feature alignment unit re-jointly encode the spliced image A1 with the text description to form a new candidate image a2, and repeat the above operations until a composite image is formed.

[0012] In summary, according to the priority selected by the user, adjust the fusion strategy. When time priority is selected, perform fast fusion; when accuracy priority is selected, perform multiple iterations for optimization, which is beneficial to more flexibly change the image synthesis method according to requirements.

[0013] Compared with the prior art, the beneficial effects of the present invention are: In the multi-modal data synthesis system based on the generative adversarial network, the dynamic hierarchical division unit recursively divides the reference image into multi-level synthesis regions based on the image feature points, and the interactive selection unit allows the user to mark the synthesis regions to be adjusted and input the text description for transforming and adjusting the synthesis regions to meet personalized needs. Then, the multi-modal feature alignment unit jointly encodes the text description and the marked multi-level synthesis regions, enhancing the consistency between the text and image features. The stage seamless stitching unit forms candidate images through the local generative adversarial network in stages, and then seamlessly stitches them with the reference image through the multi-modal fusion algorithm, ensuring the quality and consistency of the synthesized images, avoiding the limitations of global optimization, allowing for more refined local adjustments, enabling the user to mark specific regions and input descriptions to meet personalized needs. On the premise of ensuring the consistency between the text and image features and enhancing the quality and style consistency of the generation results, the user can choose time or precision priority according to needs to adapt to different usage scenarios.

[0014] In addition to the purposes, features, and advantages described above, the present invention has other purposes, features, and advantages. The present invention will be further described in detail below with reference to the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is the overall structural principle block diagram of the present invention; Figure 2 is the principle block diagram of the multi-modal feature alignment unit of the present invention; Figure 3 is the principle block diagram of the precision feedback module of the present invention.

[0016] The meanings of the various reference numerals in the figure are as follows: 100, dynamic hierarchical division unit; 110, image feature extraction module; 120, region division module; 200, interactive selection unit; 210, interaction feedback module; 220, text input module; 300, multi-modal feature alignment unit; 400, stage seamless stitching unit; 410, candidate image formation module; 420, multi-modal fusion module. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0018] Please refer to Figures 1 - 3As shown, this embodiment provides a multi-modal data synthesis system based on a generative adversarial network, including a dynamic hierarchical division unit 100, an interactive selection unit 200, a multi-modal feature alignment unit 300, and a phase seamless stitching unit 400. Therefore, to avoid the limitations of global optimization and allow for more refined local adjustments, users can mark specific regions and input descriptions to meet personalized needs. On the premise of ensuring the consistency of text and image features and enhancing the quality and style consistency of the generated results, users can select time or precision priority according to their needs to adapt to different usage scenarios. This embodiment generally includes the following steps; Step 1: The dynamic hierarchical division unit 100 performs recursive regional division on the reference image based on image feature points to form multi-level synthesis regions, improving the accuracy of regional positioning, reducing the offset of synthesized content, enabling subsequent phased generation and stitching, avoiding the limitations of global optimization, and enhancing the fineness of local adjustments; Moreover, the dynamic hierarchical division unit 100 includes an image feature extraction module 110 and a regional division module 120; The image feature extraction module 110 is used to extract the feature points of the reference image by using Canny edge detection and Harris corner detection. Among them, Canny edge detection smooths the image through Gaussian filtering to reduce noise, calculates the x and y components of the gradient using the Sobel operator, calculates the gradient magnitude and direction, retains the pixels of the local maximum, sets the rest to 0, determines the edges and potential edges, connects the potential edges to form complete edges, and then Harris corner detection extracts corner points. First, calculate the x and y gradients of the image, construct the structure tensor, weight the structure tensor using a Gaussian filter, and calculate the corner response function , where, and are the eigenvalues of the structure tensor, is a constant (usually taken as 0.06), retains the points with response values greater than the threshold as corner points, realizes the extraction of edges and corner points in the reference image. Edge points and corner points are usually located in the significant regions of the image (such as object boundaries, texture changes), and these regions have high information content in the image. In the clustering process, these feature points can be used as the basis for regional division to help divide meaningful image regions, improving the accuracy of regional positioning and the quality of subsequent image generation; The region division module 120 randomly selects K initial centroids using the K-means algorithm, assigns each feature point to the nearest centroid, calculates the centroid of each cluster, and repeats the assignment and update until the centroids no longer change, thus dividing and forming multiple initial regions. In K-means clustering, the number and distribution of feature points will affect the selection of the initial clustering center, thereby affecting the final clustering result. If there are too many feature points, the clustering result may be too refined; if there are too few feature points, the clustering result may be too rough. Therefore, when the image feature extraction module 110 extracts feature points, a feature point threshold is set according to the required accuracy of the initial region until the initial region formed by clustering can meet the required accuracy requirements, improving the accuracy of clustering. Each initial region is further subjected to feature point detection and division according to the hierarchy to form sub-regions until the set hierarchy is met. By referring to the figure according to the near recursion: initial reference figure → initial region → sub-region (sub-region 1 → sub-region 2 →... → sub-region n), where n represents the set hierarchy, the reference figure is continuously refined, enabling the staff to more accurately lock the region according to the position to be changed, avoiding inaccuracies caused by automatic text recognition of the changed region, and improving the accuracy.

[0019] Step 2: The interactive selection unit 200 allows the user to mark the composite region to be adjusted and input a text description for converting and adjusting the composite region, meeting personalized needs, where: Then, the interactive selection unit 200 includes an interaction feedback module 210 and a text input module 220; The interaction feedback module 210 is used to receive the hierarchical regions output by the region division module 120 and display them on the interactive interface. The hierarchical regions include the initial reference figure, the initial region, and the sub-regions. When the current hierarchical region is not the region required by the user, the user is allowed to set the hierarchy until a marking signal is triggered. By setting a bounding box on the interactive interface, the user can draw a rectangular box by dragging the mouse or a touch device to cover the region to be adjusted, or click on a key point within the region. The key points include hierarchy increase, hierarchy decrease, select region, and mark region, etc. When clicking on hierarchy increase, the image can be recursively advanced (i.e., the figure becomes more refined); when clicking on hierarchy decrease, the image can be recursively retreated (i.e., the figure becomes more holistic); when clicking on select region, the image currently displayed on the interactive interface can be selected to determine whether to increase or decrease the hierarchy. After clicking on select region and then clicking on mark region again, the marking of this hierarchical region can be triggered. Note that the user can mark multiple hierarchical regions to form multiple marking signals; After the text input module 220 receives the marking signal, it types in text descriptions for the hierarchical regions corresponding to multiple marking signals to guide how to adjust the hierarchical regions of the markings. For example: "Change the background within the region to blue", "Add a flower within the region", enabling the user to independently input how to adjust for each hierarchical region, which is beneficial for more precise localized adjustment.

[0020] Considering that when the user selects a region to change, if only some regions are changed, but the system generates multiple initial regions and sub-regions of the initial reference map simultaneously during operation, occupying a large operating intensity and even causing the regions that really need to be changed not to be preferentially divided. Therefore, the interaction feedback module 210 further includes a sketch adjustment module. The sketch adjustment module is used to provide a freehand drawing tool on the interactive interface, allowing the user to outline the custom regions that need to be adjusted in the initial reference map through the freehand drawing tool, enabling the image feature extraction module 110 to perceive the custom regions as the actual reference map for feature point extraction. Among them, if the initial regions are displayed on the interactive interface, they can also be outlined and selected through the freehand drawing tool, which is beneficial to avoid the large intensity caused by the simultaneous loading or operation of multiple initial regions and is beneficial for more targeted subsequent marking and improving the operation efficiency.

[0021] Step 3: The multi-modal feature alignment unit 300 is used to jointly encode the text description and the multi-order synthesis region of the marking, enhancing the consistency of the text and image features, and generating a result with the same style as the reference map. When the multi-modal feature alignment unit 300 performs joint encoding, it includes the following steps: Text embedding: Through a large amount of capturing the semantic information of the text for word segmentation processing, and encoding the text to generate a text embedding vector, using a pre-trained language model to generate the mapping relationship between the semantic information and the text embedding vector, and inputting the semantic information captured by the text description typed by the text input module 220 into the pre-trained language model to output the text embedding vector; Image feature embedding: The image feature extraction module 110 perceives the image features of multiple hierarchical regions to form the feature vector of each hierarchical region, and maps it to the same semantic space as the text embedding through a fully connected layer; Joint encoding: Perform multi-head self-attention processing on the text embedding vector and the image feature vector respectively to enhance their respective semantic representations, where: Text self-attention = , is the scaling factor, used to prevent the dot product result from being too large. , , are respectively the query, key, and value matrices of the text embedding; Image self-attention = , is the scaling factor, which is used to prevent the dot product result from being too large. , , are respectively the query, key, and value matrices of the image feature vectors. When calculating the similarity between the query and the key, an attention weight matrix is generated through the dot product and the Softmax function. At the same time, in order to capture semantic information at different levels, the multi-head attention mechanism is usually adopted to achieve deep alignment of text and image features at the semantic level, ensuring that the generated image is consistent with the text description. The multi-head attention mechanism can capture semantic information at different levels, support local adjustment of different regions of the image, improve the fineness of the generated image, and through the cross-attention mechanism, model the correlation between the text and image features corresponding to different hierarchical regions respectively, and fuse the attention outputs of the text and the image to generate the final joint encoding vector. Each hierarchical region corresponds to a joint encoding vector. Through the multi-head self-attention and cross-attention mechanisms, the model can simultaneously capture the local details and global semantics of the text and the image, achieve efficient cross-modal alignment, and can refine the correlation between the text and image features layer by layer to generate a more accurate joint encoding. At the same time, the result of the joint encoding can guide the generative adversarial network (GAN) to adjust specific regions to meet the personalized needs of users.

[0022] Step 4: Make the local generative adversarial network of the stage seamless splicing unit 400 generate candidate images in stages, and then seamlessly splice them with the reference image through the multi-modal fusion algorithm, ensuring the quality and consistency of the synthesized image; The stage seamless splicing unit 400 includes a candidate image formation module 410 and a multi-modal fusion module 420; The candidate image formation module 410 is used to arrange the hierarchical regions in the order from local to global, and sequentially receive the joint encoding vectors output by the multi-modal feature alignment unit 300 to form a joint encoding vector sequence, and generate high-quality candidate images in turn according to the generator of the local generative adversarial network, only generating images of specific hierarchical regions instead of the entire image, which helps to improve the flexibility and fineness of generation, and the discriminator distinguishes between the generated image and the real image; The multi-modal fusion module 420 is used to execute the multi-modal fusion algorithm to ensure that the text embedding and image feature embedding corresponding to the candidate image are mapped to the same space to form a fusion feature, ensuring that the two can be effectively combined, align the feature points of the generated candidate image with the reference image, and apply smoothing processing in the splicing area to ensure natural transition. The multi-modal fusion algorithm ensures that the generated image has the same style as the reference image and improves the quality of the synthesized image; And, on this basis, on the one hand, the discriminator evaluates the quality of the local image at each generation stage to ensure that each generated part conforms to the style and quality standards of the reference image. Through the feedback mechanism of the discriminator, it helps the generator gradually optimize the local image, avoid generating overly single or similar images, and maintain the diversity and richness of the images. On the other hand, the discriminator checks the spliced image to ensure that the transition between parts is natural and there are no obvious splicing marks. By identifying the inconsistencies at the splicing points, the discriminator guides the generator to make adjustments to achieve seamless splicing. This is conducive to ensuring that the style of the local image conforms to the overall reference image at each generation stage and avoiding style mismatches during the splicing process.

[0023] To cope with different scenarios, the multi-modal fusion module 420 further includes an accuracy feedback module. The accuracy feedback module is used to establish a priority interaction key between accuracy and time, allowing the user to select accuracy priority or time priority, and trigger the seamless splicing posture of the reference image, including the following postures: Posture 1: When receiving the time priority signal, map multiple candidate images to the same space and synchronously fuse them to form a complete composite image; Posture 2: When receiving the accuracy priority signal, perform staged smooth splicing on the candidate images arranged in the order from local to whole, and after each splicing, make the multi-modal feature alignment unit 300 re-union encode until a composite image is formed.

[0024] Specifically, synchronously fusing to form a complete composite image includes the following steps: After the smooth splicing of multiple candidate images A at the lowest level a is completed to form the candidate image a1 at level a - 1, continue to perform smooth splicing on multiple candidate images a1 at level a - 1 until a composite image is formed. In the case of time priority, in order to meet the requirement of quickly forming an image, directly fuse multiple candidate images. This process can directly perform smooth splicing according to the existing multiple candidate images, saving running time and only requiring one batch of joint encoding; In addition, performing staged smooth splicing on the candidate images arranged in the order from local to whole, and after each splicing, making the multi-modal feature alignment unit 300 re-union encode until a composite image is formed, includes the following steps: After the smooth stitching of multiple candidate images A under the lowest level a is completed to form the stitched image A1 at level a-1, the multi-modal feature alignment unit 300 re-performs joint encoding on the stitched image A1 and the text description to form a new candidate image a1. Then, continue to perform smooth stitching on multiple candidate images a1 under level a-1. Again, the multi-modal feature alignment unit 300 re-performs joint encoding on the stitched image A1 and the text description to form a new candidate image a2. Repeat the above operations until a composite image is formed. In this process, after each stitched image A1 is formed, joint encoding is re-performed, so that the text description is remapped to the same space as the features of the newly formed stitched image and the text embedding vector again. This is equivalent to forming the image of each level in a staged manner, which is beneficial to making the candidate image more conform to the text embedding vector. Although the running time will increase, it makes the mapping between the text and the image more accurate and ensures that the style of the local image conforms to the overall reference image.

[0025] In summary, as Figure 3 shown, according to the priority (time or accuracy) selected by the user, adjust the fusion strategy. When time is prioritized, perform fast fusion; when accuracy is prioritized, perform multiple iterative optimizations, which is beneficial to more flexibly changing the image synthesis method according to requirements.

[0026] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A multimodal data synthesis system based on generative adversarial networks, characterized by: It includes a dynamic hierarchical division unit (100), an interactive selection unit (200), a multimodal feature alignment unit (300) and a stage seamless splicing unit (400); The dynamic hierarchical division unit (100) recursively divides the reference image into regions based on image feature points to form multi-order synthesis regions, and the interactive selection unit (200) allows the user to mark the synthesis region to be adjusted and input a text description for converting and adjusting the synthesis region, wherein: The multimodal feature alignment unit (300) is used to jointly encode the text description and the marked multi-order synthesis area, so that the stage seamless splicing unit (400) forms a candidate image through a staged local generative adversarial network, and then seamlessly splices it with the reference image through a multimodal fusion algorithm.

2. The multimodal data synthesis system based on generative adversarial network according to claim 1, characterized in that: The dynamic hierarchical division unit (100) comprises an image feature extraction module (110) and a region division module (120); The image feature extraction module (110) is used to extract feature points of the reference image using Canny edge detection and Harris corner point detection; The region division module (120) uses the K-means algorithm to randomly select K initial centroids, assign each feature point to the nearest centroid, calculate the centroid of each cluster, repeat the assignment and update until the centroid no longer changes to form multiple initial regions, and further perform feature point detection and division on each initial region according to the hierarchy to form sub-regions until the set hierarchy is met.

3. The multimodal data synthesis system based on generative adversarial network according to claim 2, characterized in that: The interactive selection unit (200) comprises an interactive feedback module (210) and a text input module (220); The interactive feedback module (210) is used to receive the hierarchical area output by the area division module (120) and display it on an interactive interface, wherein the hierarchical area includes an initial reference image, an initial area and a sub-area, and when the current hierarchical area is not an area required by the user, the user is allowed to set the level until a marking signal is triggered; After receiving the marking signal, the text input module (220) enters a text description for the hierarchical areas corresponding to the multiple marking signals to guide how to adjust the marked hierarchical areas.

4. The multimodal data synthesis system based on generative adversarial network according to claim 3, characterized in that: The interactive feedback module (210) further includes a delineation adjustment module, which is used to provide a hand-drawing tool on an interactive interface, allowing a user to use the hand-drawing tool to outline a custom area that needs to be adjusted in the initial reference image, so that the image feature extraction module (110) perceives the custom area as an actual reference image to extract feature points.

5. The multimodal data synthesis system based on generative adversarial network according to claim 4, characterized in that: When the multimodal feature alignment unit (300) performs joint encoding, the following steps are included: Text embedding: performing word segmentation processing by capturing a large amount of semantic information of the text, encoding the text, generating a text embedding vector, using a pre-trained language model to generate a mapping relationship between the semantic information and the text embedding vector, recognizing the semantic information captured by the text description typed by the text input module (220) and inputting it into the pre-trained language model, and outputting a text embedding vector; Image feature embedding: the image features of multiple hierarchical regions are perceived by the image feature extraction module (110) to form a feature vector for each hierarchical region, and mapped to the same semantic space as text embedding through a fully connected layer; Joint encoding: Multi-head self-attention processing is performed on the text embedding vector and image feature vector respectively to enhance their respective semantic representations. Through the cross-attention mechanism, the correlation between the text and image features corresponding to multiple hierarchical regions is modeled, and the attention outputs of the text and image are fused to generate the final joint encoding vector, with one joint encoding vector corresponding to each hierarchical region.

6. The multimodal data synthesis system based on generative adversarial network according to claim 5, characterized in that: The stage seamless splicing unit (400) comprises a candidate image forming module (410) and a multimodal fusion module (420); The candidate image forming module (410) is used to arrange hierarchical regions in order from local to global, and sequentially receive the joint coding vectors output by the multimodal feature alignment unit (300) to form a joint coding vector sequence, sequentially generate high-quality candidate images according to the generator of the local generative adversarial network, and the discriminator distinguishes between the generated image and the real image; The multimodal fusion module (420) is used to execute a multimodal fusion algorithm, ensure that the text embedding and image feature embedding corresponding to the candidate image are mapped to the same space to form a fusion feature, align the generated candidate image with the feature points of the reference image, and apply smoothing processing to the spliced ​​area.

7. The multimodal data synthesis system based on generative adversarial network according to claim 6, characterized in that: The multimodal fusion module (420) further comprises an accuracy feedback module, wherein the accuracy feedback module is used to establish an accuracy and time priority interaction key, allowing a user to select accuracy priority or time priority, and triggering a reference image seamless splicing gesture, including the following gestures: Attitude 1, receiving the time priority signal, multiple candidate images are mapped to the same space and synchronously fused to form a complete composite image; Posture 2: When receiving the accuracy priority signal, the candidate images arranged in sequence from local to complete are smoothly spliced ​​in stages, and after each splicing, the multimodal feature alignment unit (300) is re-jointly encoded until a synthetic image is formed.

8. The multimodal data synthesis system based on generative adversarial network according to claim 7, characterized in that: The synchronous fusion forms a complete synthetic image, including the following steps: When multiple candidate images A at the lowest level a are smoothly spliced ​​to form a candidate image a1 at level a-1, the multiple candidate images a1 at level a-1 are continuously smoothly spliced ​​until a composite image is formed.

9. The multimodal data synthesis system based on generative adversarial network according to claim 8, characterized in that: The candidate images arranged in sequence from local to complete are smoothly spliced ​​in stages, and after each splicing, the multimodal feature alignment unit (300) is re-jointly encoded until a synthetic image is formed, comprising the following steps: When the multiple candidate images A at the lowest level a are smoothly spliced ​​to form a spliced ​​image A1 at level a-1, the multimodal feature alignment unit (300) re-encodes the spliced ​​image A1 and the text description to form a new candidate image a1, and continues to smoothly splice the multiple candidate images a1 at level a-1. The multimodal feature alignment unit (300) again re-encodes the spliced ​​image A1 and the text description to form a new candidate image a2, and repeats the above operation until a composite image is formed.

Citation Information

Patent Citations

  • Electronic file intelligent management method and system based on AI

    CN119226234A

  • Text mining data query method and system based on cross-modal similarity

    CN119311854A

  • Multi-modal analysis image identification system and method

    CN119832570A

  • Text Editing of Digital Images

    US20220130078A1

  • Text-to-image generation method and system based on local detail editing

    WO2024130751A1