Multi-modal Data Synthesis System Based on Generative Adversarial Network
Through dynamic hierarchical division, interactive selection and multimodal feature alignment methods, the problems of regional positioning blur and global optimization limitations in the multimodal data synthesis system are solved, and high-quality and personalized image synthesis results are achieved.
Patent Information
- Application Number
- CN202510542303.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The existing multimodal data synthesis system based on generative adversarial networks has problems with regional positioning blur, global optimization limitations and multimodal fusion bias, resulting in synthesis content offset, difficulty in fine adjustment, and insufficient alignment of text and image features.
The dynamic hierarchical division unit is used for recursive region division, combined with interactive selection unit and multimodal feature alignment unit, and joint encoding of text and image features is performed through multi-head self-attention and cross-attention mechanisms, and generated in stages and seamlessly spliced to meet personalized needs.
It improves the accuracy of regional positioning and the fineness of local adjustments, ensures that the generated results are consistent with the reference picture style, and supports users to choose time or accuracy based on their needs, and adapt to different usage scenarios.
Smart Images

Figure CN120070207B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and more specifically, to a multi-modal data synthesis system based on a generative adversarial network. Background Art
[0002] AI image synthesis technology refers to the technology of using artificial intelligence algorithms, especially machine learning and deep learning methods, to generate, modify or synthesize images, enabling computers to understand and generate complex visual content, and realizing various applications from simple image editing tasks to complex scene creation. Common AI image synthesis technologies include, but are not limited to, generative adversarial networks, which consist of a generative network and a discriminative network. Through the game process between the two, the generative network can learn how to generate realistic or appropriate images. The existing technology has been applied in many fields, such as super-resolution image generation, image conversion, and face generation, etc.
[0003] Currently, when users perform AI image synthesis, they usually send a reference image to the AI program and input text. The generative network converts the text description into an image and synthesizes it at the corresponding position of the reference image. After the discriminative network determines, the final image is automatically generated. However, the current multi-modal data synthesis system based on generative adversarial networks has the following problems:
[0004] Fuzzy region positioning: It is unable to accurately identify the image region corresponding to the user's text description, resulting in the deviation of the synthesized content.
[0005] Limited global optimization: Traditional methods perform unified feature extraction on the entire image, making it difficult to finely adjust local regions.
[0006] Multi-modal fusion deviation: The alignment of text and image features is insufficient, and the generated result is inconsistent with the style of the reference image. In view of this, we propose a multi-modal data synthesis system based on generative adversarial networks. Summary of the Invention
[0007] The purpose of the present invention is to provide a multi-modal data synthesis system based on generative adversarial networks to solve the problem that the accuracy of AI in recognizing text and pictures is limited in the above background art, especially when making a separate adjustment to a certain area of the image, the synthesized image is deviated due to the inability to accurately identify the position and content corresponding to the text.
[0008] To achieve the above purpose, the present invention provides a multi-modal data synthesis system based on generative adversarial networks, including a dynamic hierarchical division unit, an interactive selection unit, a multi-modal feature alignment unit, and a stage seamless splicing unit;
[0009] The dynamic hierarchical partitioning unit recursively partitions the reference image based on image feature points to form multi-level composite regions, improving the accuracy of region localization, reducing the offset of composite content, enabling subsequent phased generation and stitching, avoiding global optimization limitations, enhancing the fineness of local adjustment, and allowing the interactive selection unit to permit the user to mark the composite regions to be adjusted and input text descriptions for transforming and adjusting the composite regions, meeting personalized requirements, where:
[0010] The multi-modal feature alignment unit is used to jointly encode the text description and the marked multi-level composite regions, enhancing the consistency between text and image features, with the generated result being consistent with the style of the reference image. The stage seamless stitching unit forms candidate images through the local generative adversarial network in stages, and then seamlessly stitches them with the reference image through the multi-modal fusion algorithm, ensuring the quality and consistency of the composite image.
[0011] As a further improvement of this technical solution, the dynamic hierarchical partitioning unit includes an image feature extraction module and a region partitioning module;
[0012] The image feature extraction module is used to extract the feature points of the reference image by using Canny edge detection and Harris corner detection. During the clustering process, these feature points can be used as the basis for region partitioning, helping to divide meaningful image regions, improving the accuracy of region localization and the quality of subsequent image generation;
[0013] The region partitioning module randomly selects K initial centroids using the K-means algorithm, assigns each feature point to the nearest centroid, calculates the centroid of each cluster, repeats the assignment and update until the centroids no longer change, and partitions to form multiple initial regions. Further feature point detection and partitioning are performed on each initial region according to the hierarchy to form sub-regions until the set hierarchy is satisfied, realizing the continuous refinement of the reference image, enabling the staff to more accurately lock the region according to the position to be changed, avoiding inaccuracies caused by automatic text recognition of the changed region, and improving the accuracy.
[0014] As a further improvement of this technical solution, the interactive selection unit includes an interaction feedback module and a text input module;
[0015] The interaction feedback module is used to receive the hierarchical regions output by the region partitioning module and display them on the interactive interface. The hierarchical regions include the initial reference image, initial regions, and sub-regions. When the current hierarchical region is not the region required by the user, the user is allowed to set the hierarchy until a marking signal is triggered, and the user can mark multiple hierarchical regions to form multiple marking signals;
[0016] After receiving the marking signals, the text input module types in text descriptions for the hierarchical regions corresponding to multiple marking signals to guide how to adjust the hierarchical regions of the markings, enabling the user to independently input how to make adjustments for each hierarchical region, which is conducive to more precise localized adjustments.
[0017] As a further improvement of this technical solution, the interaction feedback module further includes a sketch adjustment module. The sketch adjustment module is used to provide a hand-drawing tool on the interactive interface, allowing the user to outline a custom region that needs to be adjusted in the initial reference diagram through the hand-drawing tool, enabling the image feature extraction module to perceive the custom region as the actual reference diagram for feature point extraction. This is conducive to avoiding the high intensity caused by the simultaneous loading or operation of multiple initial regions, facilitating more targeted subsequent marking, and improving the operation efficiency.
[0018] As a further improvement of this technical solution, when the multi-modal feature alignment unit performs joint encoding, it includes the following steps:
[0019] Text embedding: Through a large amount of semantic information capture of the text, word segmentation processing is carried out, and the text is encoded to generate a text embedding vector. A pre-trained language model is used to generate the mapping relationship between the semantic information and the text embedding vector. The semantic information captured by the text description typed in by the text input module is identified and input into the pre-trained language model to output the text embedding vector;
[0020] Image feature embedding: The image feature extraction module perceives the image features of multiple hierarchical regions to form the feature vector of each hierarchical region, and maps it to the same semantic space as the text embedding through a fully connected layer;
[0021] Joint encoding: Perform multi-head self-attention processing on the text embedding vector and the image feature vector respectively to enhance their respective semantic representations. Through the cross-attention mechanism, model the correlation between the text and image features corresponding to multiple hierarchical regions respectively, fuse the attention outputs of the text and the image to generate the final joint encoding vector. Each hierarchical region corresponds to a joint encoding vector. Through the multi-head self-attention and cross-attention mechanisms, the model can simultaneously capture the local details and global semantics of the text and the image, achieve efficient cross-modal alignment, and can gradually refine the correlation between the text and image features to generate more accurate joint encoding. At the same time, the result of the joint encoding can guide the generative adversarial network to make adjustments to specific regions to meet the personalized needs of users.
[0022] As a further improvement of this technical solution, the stage seamless splicing unit includes a candidate image formation module and a multi-modal fusion module;
[0023] The candidate image formation module is used to arrange hierarchical regions in the order from local to global, and sequentially receive the jointly encoded vectors output by the multimodal feature alignment unit to form a sequence of jointly encoded vectors. According to the generator of the local generative adversarial network, high-quality candidate images are generated in sequence, and only images of specific hierarchical regions are generated instead of the entire image, which helps to improve the flexibility and fineness of generation. The discriminator distinguishes between the generated images and the real images;
[0024] The multimodal fusion module is used to execute a multimodal fusion algorithm to ensure that the text embedding and the image feature embedding corresponding to the candidate image are mapped to the same space to form a fused feature, ensuring that the two can be effectively combined. Align the feature points of the generated candidate image with those of the reference image, and apply smoothing processing in the splicing area to ensure natural transition.
[0025] As a further improvement of this technical solution, the multimodal fusion module further includes an accuracy feedback module. The accuracy feedback module is used to establish a priority interaction key between accuracy and time, allowing the user to select accuracy priority or time priority, and trigger the seamless splicing posture of the reference image, including the following postures:
[0026] Posture 1: Receive the time priority signal, then map multiple candidate images to the same space and synchronously fuse them to form a complete composite image;
[0027] Posture 2: Receive the accuracy priority signal, then perform staged smooth splicing on the candidate images arranged in the order from local to global, and after each splicing, make the multimodal feature alignment unit perform joint encoding again until a composite image is formed.
[0028] As a further improvement of this technical solution, the synchronous fusion to form a complete composite image includes the following steps:
[0029] After the smooth splicing of multiple candidate images A at the lowest level a is completed to form a candidate image a1 at level a - 1, continue to perform smooth splicing on multiple candidate images a1 at level a - 1 until a composite image is formed;
[0030] The staged smooth splicing of the candidate images arranged in the order from local to global, and after each splicing, make the multimodal feature alignment unit perform joint encoding again until a composite image is formed, includes the following steps:
[0031] After the smooth stitching of multiple candidate images A under the lowest level a is completed to form the stitched image A1 at level a - 1, the multimodal feature alignment unit jointly encodes the stitched image A1 and the text description again to form a new candidate image a1, and continues to perform smooth stitching on multiple candidate images a1 under level a - 1. Again, the multimodal feature alignment unit jointly encodes the stitched image A1 and the text description to form a new candidate image a2, and repeats the above operations until a composite image is formed.
[0032] In summary, according to the priority selected by the user, the fusion strategy is adjusted. When time is prioritized, rapid fusion is performed; when accuracy is prioritized, multiple iterations of optimization are carried out, which is beneficial for more flexible changes in the image synthesis method according to requirements.
[0033] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0034] In this multi-modal data synthesis system based on a generative adversarial network, the dynamic hierarchical division unit recursively divides the reference image into multi-level synthesis regions based on image feature points, and the interactive selection unit allows the user to mark the synthesis regions that need to be adjusted and input text descriptions for transforming and adjusting the synthesis regions to meet personalized needs. Then, the multimodal feature alignment unit jointly encodes the text description and the marked multi-level synthesis regions, enhancing the consistency between text and image features. The stage seamless stitching unit forms candidate images through a local generative adversarial network in stages, and then seamlessly stitches them with the reference image through a multi-modal fusion algorithm, ensuring the quality and consistency of the synthesized image, avoiding the limitations of global optimization, allowing for more refined local adjustments, enabling the user to mark specific regions and input descriptions to meet personalized needs. On the premise of enhancing the quality and style consistency of the generation result while ensuring the consistency between text and image features, the user can select time or accuracy priority according to requirements to adapt to different usage scenarios.
[0035] In addition to the purposes, features, and advantages described above, the present invention has other purposes, features, and advantages. The present invention will be further described in detail below with reference to the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is the overall structural principle block diagram of the present invention;
[0037] Figure 2 is the principle block diagram of the multimodal feature alignment unit of the present invention;
[0038] Figure 3 is the principle block diagram of the accuracy feedback module of the present invention.
[0039] The meanings of the various reference numerals in the figure are as follows:
[0040] 100, Dynamic hierarchical partitioning unit; 110, Image feature extraction module; 120, Region partitioning module;
[0041] 200, Interactive selection unit; 210, Interaction feedback module; 220, Text input module;
[0042] 300, Multimodal feature alignment unit;
[0043] 400, Phase seamless splicing unit; 410, Candidate image formation module; 420, Multimodal fusion module. Detailed implementation manner
[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0045] Please refer to Figures 1 - 3 As shown, this embodiment provides a multimodal data synthesis system based on a generative adversarial network, including a dynamic hierarchical partitioning unit 100, an interactive selection unit 200, a multimodal feature alignment unit 300, and a phase seamless splicing unit 400. Therefore, in order to avoid the limitations of global optimization and allow for more refined local adjustments, users can mark specific regions and input descriptions to meet personalized needs. On the premise of ensuring the consistency of text and image features and enhancing the quality and style consistency of the generated results, users can select time or accuracy priority according to their needs to adapt to different usage scenarios. This embodiment generally includes the following steps;
[0046] Step 1: The dynamic hierarchical partitioning unit 100 recursively partitions the reference image based on image feature points to form multi-level synthesis regions, improving the accuracy of region positioning, reducing the offset of synthesized content, realizing subsequent staged generation and splicing, avoiding the limitations of global optimization, and improving the fineness of local adjustments;
[0047] Moreover, the dynamic hierarchical partitioning unit 100 includes an image feature extraction module 110 and a region partitioning module 120;
[0048] The image feature extraction module 110 is used to extract feature points of the reference image by using Canny edge detection and Harris corner detection. Among them, Canny edge detection smooths the image through Gaussian filtering to reduce noise, calculates the x and y components of the gradient using the Sobel operator, and calculates the gradient magnitude and direction. The pixels of the local maximum are retained, and the rest are set to 0 to determine the edges and potential edges, and the potential edges are connected to form complete edges. Then, Harris corner detection extracts corner points. First, calculate the x and y gradients of the image, construct the structure tensor, weight the structure tensor using a Gaussian filter, and calculate the corner response function , where and are the eigenvalues of the structure tensor, is a constant (usually taken as 0.06). The points with response values greater than the threshold are retained as corner points, realizing the extraction of edges and corner points in the reference image. Edge points and corner points are usually located in the significant regions of the image (such as object boundaries, texture change points). These regions have higher information content in the image, which is beneficial in the clustering process. These feature points can be used as the basis for region division, helping to divide meaningful image regions and improving the accuracy of region positioning and the quality of subsequent image generation;
[0049] The region division module 120 uses the K-means algorithm to randomly select K initial centroids, assigns each feature point to the nearest centroid, calculates the centroid of each cluster, and repeats the assignment and update until the centroids no longer change to divide and form multiple initial regions. In K-means clustering, the number and distribution of feature points will affect the selection of the initial clustering center, thus affecting the final clustering result. If there are too many feature points, it may lead to an overly refined clustering result; if there are too few feature points, it may lead to an overly rough clustering result. Therefore, when the image feature extraction module 110 extracts feature points, a feature point threshold is set according to the required accuracy of the initial region until the initial regions formed by clustering can meet the required accuracy requirements, improving the accuracy of clustering. Each initial region is further subjected to feature point detection and division according to the hierarchy to form sub-regions until the set hierarchy is satisfied. Through the reference image in a near-recursive manner: initial reference image → initial region → sub-region (sub-region 1 → sub-region 2 →... → sub-region n), where n represents the set hierarchy, realizing the continuous refinement of the reference image, enabling the staff to more accurately lock the region according to the position to be changed, avoiding inaccuracies caused by automatic text recognition of the changed region, and improving the accuracy.
[0050] Step 2: Make the interactive selection unit 200 allow the user to mark the composite region to be adjusted and input a text description for converting and adjusting the composite region to meet personalized needs, where:
[0051] Then, the interactive selection unit 200 includes an interactive feedback module 210 and a text input module 220;
[0052] The interactive feedback module 210 is used to receive the hierarchical regions output by the region division module 120 and display them on the interactive interface. The hierarchical regions include the initial reference diagram, the initial regions, and the sub-regions. When the current hierarchical region is not the region required by the user, the user is allowed to set the hierarchy until a marking signal is triggered. By setting a bounding box on the interactive interface, the user can draw a rectangular box by dragging the mouse or a touch device to cover the area to be adjusted, or perform point selection, that is, the user clicks on the key points within the region. The key points include hierarchy increase, hierarchy decrease, select region, and mark region, etc. When clicking on hierarchy increase, the image can be recursively forward (that is, making the diagram more refined), and when clicking on hierarchy decrease, the image can be recursively backward (that is, making the diagram more holistic). When clicking on the select region, the image currently displayed on the interactive interface can be selected to determine whether to increase or decrease the hierarchy. After clicking on the select region and then clicking on the mark region again, the marking of this hierarchical region can be triggered. Note that the user can mark multiple hierarchical regions to form multiple marking signals;
[0053] After receiving the marking signal, the text input module 220 types in a text description for the hierarchical regions corresponding to the multiple marking signals to guide how to adjust the marked hierarchical regions. For example: "Change the background within the region to blue", "Add a flower within the region", so that the user can independently input how to adjust for each hierarchical region, which is beneficial for more precise localized adjustment.
[0054] Considering that when the user selects to change the region, if only some regions are changed, but the system generates multiple initial regions and sub-regions of the initial reference diagram simultaneously during operation, which occupies a large operating intensity and may even cause the truly required changed region not to be preferentially divided. Therefore, the interactive feedback module 210 also includes a sketch adjustment module. The sketch adjustment module is used to provide a freehand drawing tool on the interactive interface, allowing the user to outline the custom region to be adjusted in the initial reference diagram through the freehand drawing tool, enabling the image feature extraction module 110 to perceive the custom region as the actual reference diagram for feature point extraction. Among them, if the initial region is displayed on the interactive interface, it can also be outlined and selected through the freehand drawing tool, which is beneficial for avoiding the large intensity caused by the simultaneous loading or operation of multiple initial regions and is beneficial for more targeted subsequent marking and improving the operation efficiency.
[0055] Step 3: The multimodal feature alignment unit 300 is used to jointly encode the text description and the marked multi-level composite region, enhancing the consistency between the text and image features, and generating a result with the same style as the reference diagram.
[0056] When the multimodal feature alignment unit 300 performs joint encoding, it includes the following steps:
[0057] Text Embedding: Tokenize the text by capturing a large amount of semantic information of the text, encode the text, generate text embedding vectors, use a pre-trained language model to generate the mapping relationship between semantic information and text embedding vectors, identify the semantic information captured by the text description typed by the text input module 220, and input it into the pre-trained language model to output text embedding vectors;
[0058] Image Feature Embedding: The image feature extraction module 110 perceives the image features of multiple hierarchical regions to form feature vectors for each hierarchical region, and maps them to the same semantic space as the text embedding through a fully connected layer;
[0059] Joint Encoding: Perform multi-head self-attention processing on the text embedding vectors and image feature vectors respectively to enhance their respective semantic representations, where:
[0060] Text Self-Attention = , is a scaling factor used to prevent the dot product result from being too large, 、 、 are the query, key, and value matrices of the text embedding respectively;
[0061] Image Self-Attention = , is a scaling factor used to prevent the dot product result from being too large, 、 、 are the query, key, and value matrices of the image feature vector respectively. When calculating the similarity between the query and the key, an attention weight matrix is generated through the dot product and the Softmax function. At the same time, in order to capture semantic information at different levels, a multi-head attention mechanism is usually adopted to achieve deep alignment of text and image features at the semantic level, ensuring that the generated image is consistent with the text description. The multi-head attention mechanism can capture semantic information at different levels, support local adjustment of different regions of the image, improve the fineness of the generated image, and through the cross-attention mechanism, model the correlation between the text and image features corresponding to multiple hierarchical regions respectively, fuse the attention outputs of the text and the image, and generate the final joint encoding vector. Each hierarchical region corresponds to a joint encoding vector. Through the multi-head self-attention and cross-attention mechanisms, the model can simultaneously capture the local details and global semantics of the text and the image, achieve efficient cross-modal alignment, and can gradually refine the correlation between the text and image features to generate more accurate joint encoding. At the same time, the result of the joint encoding can guide the generative adversarial network (GAN) to adjust specific regions to meet the personalized needs of users.
[0062] Step 4: The phased seamless stitching unit 400 uses the local generative adversarial network in phases to form candidate images, and then seamlessly stitches them with the reference image through a multimodal fusion algorithm, ensuring the quality and consistency of the synthesized images;
[0063] The phased seamless stitching unit 400 includes a candidate image formation module 410 and a multimodal fusion module 420;
[0064] The candidate image formation module 410 is used to arrange the hierarchical regions in the order from local to global, and sequentially receive the jointly encoded vectors output by the multimodal feature alignment unit 300 to form a sequence of jointly encoded vectors. According to the generator of the local generative adversarial network, high-quality candidate images are generated in sequence, generating only the images of specific hierarchical regions instead of the entire image, which helps to improve the flexibility and fineness of generation. The discriminator distinguishes between the generated images and the real images;
[0065] The multimodal fusion module 420 is used to execute the multimodal fusion algorithm, ensuring that the text embedding and image feature embedding corresponding to the candidate image are mapped to the same space to form fusion features, ensuring that the two can be effectively combined, aligning the feature points of the generated candidate image with the reference image, and applying smoothing processing in the stitching area to ensure natural transition. The multimodal fusion algorithm ensures that the generated image is consistent with the style of the reference image and improves the quality of the synthesized image;
[0066] Moreover, on this basis, on the one hand, the discriminator evaluates the quality of the local images in each generation stage, ensuring that each generated part conforms to the style and quality standards of the reference image. Through the feedback mechanism of the discriminator, it helps the generator gradually optimize the local images, avoiding the generation of overly single or similar images and maintaining the diversity and richness of the images. On the other hand, the discriminator checks the stitched image, ensuring that the transition between parts is natural and there are no obvious stitching traces. By identifying the inconsistencies at the stitching points by the discriminator, it guides the generator to make adjustments to achieve seamless stitching, which is beneficial to ensuring that the style of the local images conforms to the overall reference image in each generation stage and avoiding style mismatches during the stitching process.
[0067] To cope with different scenarios, the multimodal fusion module 420 further includes an accuracy feedback module. The accuracy feedback module is used to establish a priority interaction key between accuracy and time, allowing the user to select accuracy priority or time priority, triggering the seamless stitching posture of the reference image, including the following postures:
[0068] Posture 1: When receiving a time priority signal, map multiple candidate images to the same space and synchronously fuse them to form a complete synthesized image;
[0069] Pose 2: Receive the precision - priority signal, then perform staged smooth stitching on the candidate images arranged in the order from local to global. After each stitching, make the multimodal feature alignment unit 300 re - jointly encode until a composite image is formed.
[0070] Specifically, synchronous fusion to form a complete composite image includes the following steps:
[0071] After the smooth stitching of multiple candidate images A at the lowest level a is completed to form the candidate image a1 at level a - 1, continue to perform smooth stitching on multiple candidate images a1 at level a - 1 until a composite image is formed. In the case of time priority, in order to satisfy the rapid formation of the image, directly fuse multiple candidate images. This process can directly perform smooth stitching according to the existing multiple candidate images, saving the running time and only requiring a batch of joint encoding.
[0072] In addition, perform staged smooth stitching on the candidate images arranged in the order from local to global, and after each stitching, make the multimodal feature alignment unit 300 re - jointly encode until a composite image is formed, including the following steps:
[0073] After the smooth stitching of multiple candidate images A at the lowest level a is completed to form the stitched image A1 at level a - 1, make the multimodal feature alignment unit 300 re - jointly encode the stitched image A1 with the text description to form a new candidate image a1. Continue to perform smooth stitching on multiple candidate images a1 at level a - 1. Again, make the multimodal feature alignment unit 300 re - jointly encode the stitched image A1 with the text description to form a new candidate image a2. Repeat the above operations until a composite image is formed. In this process, after each formation of the stitched image A1, re - joint encoding is performed, so that the text description is remapped to the same space again with the features of the newly formed stitched image and the text embedding vector, which is equivalent to forming the image of each level in stages, facilitating the candidate image to better fit the text embedding vector. Although the running time will increase, it makes the mapping between the text and the image more accurate and ensures that the style of the local image conforms to the overall reference image.
[0074] In summary, as Figure 3 shown, adjust the fusion strategy according to the priority (time or precision) selected by the user. When time is prioritized, perform rapid fusion; when precision is prioritized, perform multiple iterations of optimization, which is beneficial to more flexibly change the image synthesis method according to requirements.
[0075] The basic principles, main features and advantages of the present invention have been shown and described above. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.
Claims
1. A multi-modal data synthesis system based on a generative adversarial network, characterized in that: It includes a dynamic hierarchical partitioning unit (100), an interactive selection unit (200), a multimodal feature alignment unit (300), and a stage seamless splicing unit (400); The dynamic hierarchical partitioning unit (100) recursively partitions the reference image based on image feature points to form multi-level composite regions, and the interactive selection unit (200) allows the user to mark the composite regions to be adjusted and input text descriptions for converting and adjusting the composite regions, where: The multimodal feature alignment unit (300) is used to jointly encode the text description and the marked multi-level composite regions, enabling the stage seamless splicing unit (400) to form candidate images through a stage-based local generative adversarial network, and then seamlessly splicing with the reference image through a multimodal fusion algorithm; The dynamic hierarchical partitioning unit (100) includes an image feature extraction module (110) and a region partitioning module (120); the image feature extraction module (110) is used to extract the feature points of the reference image by using Canny edge detection and Harris corner detection; the region partitioning module (120) randomly selects K initial centroids using the K-means algorithm, assigns each feature point to the nearest centroid, calculates the centroid of each cluster, repeats the assignment and update until the centroids no longer change to partition and form multiple initial regions, and further performs feature point detection and partitioning on each initial region according to the hierarchy to form sub-regions until the set hierarchy is satisfied; The interactive selection unit (200) includes an interaction feedback module (210) and a text input module (220); the interaction feedback module (210) is used to receive the hierarchical regions output by the region partitioning module (120) and display them on the interactive interface. The hierarchical regions include the initial reference image, the initial regions, and the sub-regions. When the current hierarchical region is not the region required by the user, the user is allowed to set the hierarchy until a marking signal is triggered; after receiving the marking signal, the text input module (220) types in text descriptions for the hierarchical regions corresponding to multiple marking signals to guide how to adjust the marked hierarchical regions; When the multimodal feature alignment unit (300) performs joint encoding, it includes the following steps: Text embedding: Perform word segmentation processing by capturing a large amount of semantic information of the text, encode the text to generate a text embedding vector, use a pre-trained language model to generate the mapping relationship between the semantic information and the text embedding vector, identify the semantic information captured by the text description typed by the text input module (220), input it into the pre-trained language model, and output the text embedding vector; Image feature embedding: The image feature extraction module (110) perceives the image features of multiple hierarchical regions to form the feature vector of each hierarchical region, and maps it to the same semantic space as the text embedding through a fully connected layer; Joint encoding: Perform multi-head self-attention processing on the text embedding vector and the image feature vector respectively to enhance their respective semantic representations. Through the cross-attention mechanism, model the correlation between the text and the image features corresponding to multiple hierarchical regions, fuse the attention outputs of the text and the image to generate the final joint encoding vector, and each hierarchical region corresponds to a joint encoding vector.
2. The multi-modal data synthesis system based on a generative adversarial network according to claim 1, wherein: The interaction feedback module (210) further includes a sketch adjustment module. The sketch adjustment module is used to provide a hand-drawing tool on the interactive interface, allowing the user to use the hand-drawing tool to outline the custom region that needs to be adjusted in the initial reference map, so that the image feature extraction module (110) perceives the custom region as the actual reference map for feature point extraction.
3. The multi-modal data synthesis system based on a generative adversarial network according to claim 2, wherein: The stage seamless stitching unit (400) includes a candidate image formation module (410) and a multimodal fusion module (420); The candidate image formation module (410) is used to arrange the hierarchical regions in the order from local to global, and sequentially receive the joint encoding vectors output by the multimodal feature alignment unit (300) to form a joint encoding vector sequence. According to the generator of the local generative adversarial network, generate high-quality candidate images in sequence, and the discriminator distinguishes between the generated images and the real images; The multimodal fusion module (420) is used to execute the multimodal fusion algorithm to ensure that the text embedding and the image feature embedding corresponding to the candidate image are mapped to the same space to form a fusion feature, align the feature points of the generated candidate image with the reference map, and apply smoothing processing in the stitching region.
4. The multi-modal data synthesis system based on a generative adversarial network according to claim 3, characterized in that: The multimodal fusion module (420) further includes an accuracy feedback module. The accuracy feedback module is used to establish a priority interaction key for accuracy and time, allowing the user to select accuracy priority or time priority, and trigger the seamless stitching posture of the reference map, including the following postures: Posture 1: Receive the time priority signal, then map multiple candidate images to the same space and synchronously fuse them to form a complete composite image; Posture 2: Receive the accuracy priority signal, then perform staged smooth stitching on the candidate images arranged in the order from local to global, and after each stitching, make the multimodal feature alignment unit (300) re-jointly encode until a composite image is formed.
5. The multi-modal data synthesis system based on a generative adversarial network according to claim 4, characterized in that: The synchronous fusion to form a complete composite image includes the following steps: After the smooth stitching of multiple candidate images A at the lowest level a is completed to form the candidate image a1 at level a - 1, continue to perform smooth stitching on multiple candidate images a1 at level a - 1 until a composite image is formed.
6. The multi-modal data synthesis system based on a generative adversarial network according to claim 5, characterized in that: The staged smooth stitching of the candidate images arranged in the order from local to global, and after each stitching, make the multimodal feature alignment unit (300) re-jointly encode until a composite image is formed, includes the following steps: After the smooth stitching of multiple candidate images A under the lowest level a is completed to form the stitched image A1 at level a - 1, the multi-modal feature alignment unit (300) re-performs joint encoding on the stitched image A1 and the text description to form a new candidate image a1, and continues to perform smooth stitching on multiple candidate images a1 under level a - 1. Again, the multi-modal feature alignment unit (300) re-performs joint encoding on the stitched image A1 and the text description to form a new candidate image a2, and repeats the above operations until a composite image is formed.
Citation Information
Patent Citations
Electronic file intelligent management method and system based on AI
CN119226234A
Text mining data query method and system based on cross-modal similarity
CN119311854A