Modal language model image editing technology fusing non-perpetual culture elements
By introducing a modal language model and two-way interaction module that integrates intangible cultural elements into the instruction image editing technology, the problems of multi-object editing difficulties and insufficient complex reasoning in complex image editing tasks are solved, and higher editing accuracy and cultural fidelity are achieved.
Patent Information
- Application Number
- CN202510180691.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-10-25
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-10
AI Technical Summary
In the prior art, when dealing with complex image editing tasks, especially when editing multi-object and intangible cultural heritage elements, there are problems such as multi-object editing difficulties, insufficient complex reasoning scenarios and limitations of information interaction.
A modal language model image editing technology that integrates intangible cultural elements is adopted. By selecting the LLaMA model and introducing LoRA for adaptive fine-tuning, combining the two-way interaction module and diffusion model, multiple rounds of information flow and deep interaction between the image and text features are achieved.
It significantly improves the model's understanding and reasoning ability of complex instructions, improves the accuracy of multi-object editing and the performance of complex cultural reasoning tasks, and ensures the uniqueness of cultural symbols preserved during image generation.
Smart Images

Figure CN120125946A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of instruction image editing, and particularly to a modal language model image editing technology integrating intangible cultural heritage elements. Background Art
[0002] Instruction image editing is a technology where users modify images through natural language instructions. This field has developed rapidly in recent years. To help intangible cultural heritage better inherit and develop in the digital age and empower the inheritance of intangible cultural heritage with digital intelligence technology, we propose an image editing method combining a multi-modal large language model (MLLM), which uses digital intelligence technology for value co-creation of intangible cultural heritage craftsmen's products, called "immersive experience". This method aims to enhance the understanding and reasoning abilities of existing instruction editing technologies, especially in complex instruction scenarios, such as dealing with multiple objects, multiple attributes, etc., and is ultimately applied to the digital transformation and innovative development of intangible cultural heritage represented by the "Three Knives" flower-inserting technique.
[0003] In the field of instruction image editing, existing technologies such as "InstructPix2Pix" rely on simple text encoders to parse the instructions of diffusion models. These methods work well when dealing with simple image editing instructions, but in the face of complex scenarios, especially when editing elements related to intangible cultural heritage (ICH), the following challenges exist:
[0004] Difficulty in multi-object editing: When an image contains multiple objects and the instruction specifies modifying a specific attribute (such as position, size, color) of a certain object, existing methods often have difficulty in accurately executing.
[0005] Insufficiency in complex reasoning scenarios: Some editing instructions require the model to have world knowledge (such as time, space relationships or common sense) for reasoning, and existing methods perform poorly in understanding and identifying ICH elements to be edited.
[0006] Limitations in information interaction: In traditional methods, the interaction between images and texts is mostly one-way, lacking effective two-way information flow, which limits the ability to deeply understand and edit ICH elements. Summary of the Invention
[0007] The purpose of the present invention is to provide a modal language model image editing technology integrating intangible cultural heritage elements, which has the advantage of improving the model's understanding and reasoning abilities for complex instructions, and solves the problems raised in the background art.
[0008] To achieve the above purpose, the present invention provides the following technical solution: A modal language model image editing technology integrating intangible cultural heritage elements, including:
[0009] Model Selection and Training: Select the LLaMA model as the basis and introduce LoRA for adaptive fine-tuning. In this way, the model can be adaptively adjusted while keeping the original parameters frozen, greatly improving the training efficiency and instruction understanding ability. Use training data related to intangible cultural heritage for model training.
[0010] Bidirectional Interaction Module: The bidirectional interaction module includes a feature extraction layer, an interaction layer, and a fusion layer. The feature extraction layer includes text feature extraction and image feature extraction. The text feature extraction uses a pre-trained text encoder to extract the feature vector of the text instruction; the image feature extraction uses a pre-trained image encoder to extract the feature vector of the image. The interaction layer includes text-to-image interaction and image-to-text interaction. The text-to-image interaction inputs the text feature vector as conditional information into the image generation or editing model to affect the image generation or editing process. The image-to-text interaction feeds back the image feature vector to the text processing module to adjust the understanding of the text instruction or generate an image-related text description. The fusion layer fuses the text and image feature vectors output by the interaction layer to generate a fused feature vector for subsequent image generation and editing tasks, realizing multi-round information flow between image and text features. Verify the effectiveness of the bidirectional interaction module through experiments and adjust the model parameters to optimize the performance.
[0011] Diffusion Model: The diffusion model generates images through gradual iteration. Starting from a random noise image, through the gradual guidance and generation process of the instruction, the model gradually approaches the target image. Therefore, when dealing with complex intangible cultural heritage image editing tasks, the diffusion model can ensure that the generated results have sufficient details and accuracy through its phased generation strategy. The diffusion model can gradually generate these details to keep them highly consistent with the original image and conform to the design style of cultural elements. Combined with the bidirectional interaction module, the diffusion model can not only generate high-quality images that meet the user's requirements when executing editing instructions, but also ensure the uniqueness of cultural symbols in the image generation process through interaction with text features.
[0012] Preferably, the two-way interaction module in step 2 includes: a feature extraction layer, an interaction layer, and a fusion layer. The feature extraction layer includes text feature extraction and image feature extraction. The text feature extraction uses a pre-trained text encoder to extract the feature vector of the text instruction; the image feature extraction uses a pre-trained image encoder to extract the feature vector of the image; the interaction layer includes text-to-image interaction and image-to-text interaction. The text-to-image interaction inputs the text feature vector as conditional information into the image generation or editing model to affect the image generation or editing process. The image-to-text interaction feeds back the image feature vector to the text processing module for adjusting the understanding of the text instruction or generating an image-related text description. The fusion layer fuses the text and image feature vectors output by the interaction layer to generate a fused feature vector for subsequent image generation and editing tasks, realizing multi-round information flow between image and text features. The effectiveness of the two-way interaction module is verified through experiments, and the model parameters are adjusted to optimize the performance.
[0013] Preferably, the training data related to intangible cultural heritage in step 1 includes image segmentation data, complex instruction editing data, and world knowledge reasoning data.
[0014] Preferably, the two-way interaction module further includes an attention mechanism. The attention mechanism strengthens the association between image and text features. In the text-to-image interaction, the attention mechanism is used to focus on the image regions most relevant to the text instruction; in the image-to-text interaction, the attention mechanism is used to focus on the text parts most relevant to the image features.
[0015] Preferably, the two-way interaction module further includes a multimodal fusion mechanism. The multimodal fusion mechanism effectively fuses the text and image feature vectors in the fusion layer. The multimodal fusion mechanism includes but is not limited to concatenation, weighted summation, and bilinear pooling. The two-way interaction module can combine the text and image feature vectors to form a unified multimodal feature representation. The fused multimodal feature representation is used in subsequent image generation and editing tasks. Through this multimodal feature representation, the system can simultaneously understand the text instruction and the image content, thus achieving a more accurate and user-expected image editing result.
[0016] Preferably, the specific implementation steps of the text-to-image interaction are as follows: when a user inputs a natural language instruction containing intangible cultural heritage elements, the system first uses a pre-trained text encoder to extract the feature vector of the text instruction. This feature vector is then input into the image generation or editing model as conditional information, affecting the image generation or editing process. In this way, the descriptions and instructions in the text can directly guide the image generation to ensure that the image content meets the user's expectations. The specific implementation steps of the image-to-text interaction are as follows: during the image editing process, a pre-trained image encoder is used to extract the feature vector of the image. These image feature vectors are then fed back to the text processing module to adjust the understanding of the text instruction or generate a text description related to the image. This interaction mechanism enables the system to optimize the interpretation of the text instruction according to the actual situation of the image, ensuring that the editing result is more accurate and in line with the user's intention. That is, text feature extraction encodes the natural language instruction using a pre-trained text encoder. The encoder converts the text into a series of high-dimensional feature vectors that can capture the semantic information and context relationships in the text. Image feature extraction encodes the image using a pre-trained image encoder. The encoder converts the image into a series of feature maps or feature vectors that can capture the visual information and structural information in the image.
[0017] Preferably, the diffusion model goes through an initial noise stage, a gradual denoising stage, and a result fusion stage; the result fusion is implicitly completed during the iteration of each time step. The model at each time step will perform further denoising and refinement based on the output of the previous time step. Finally, when the preset number of iterations is reached, the model will output a fully denoised and highly refined data sample as the final result. The generation strategy of gradual denoising decomposes the complex generation task into iterative subtasks at multiple time steps. During the training stage, the model learns how to reverse the forward diffusion process, that is, to recover the original data from the noisy data. During the inference stage, the model randomly samples a noise vector from the standard Gaussian distribution and then gradually removes the noise through the reverse diffusion process to finally generate a clear image or other types of data samples.
[0018] Preferably, the data augmentation strategy includes:
[0019] Introduce various data sources related to intangible cultural heritage: Image segmentation data: This type of data contains fine-grained annotations of intangible cultural heritage elements, which helps the model better identify and understand these elements during the learning process. Complex instruction editing data: By providing editing data containing complex instructions, the model can learn how to precisely edit the intangible cultural heritage elements in the image under the guidance of natural language instructions. World knowledge reasoning data: This type of data contains rich cultural background knowledge and common sense information, which helps the model better understand the implicit intentions in the user's instructions during the reasoning process;
[0020] Use diverse data augmentation techniques: Image augmentation: Perform operations such as rotation, scaling, cropping, and color transformation on the image to increase the diversity and robustness of the image. These operations can simulate different shooting conditions and perspectives, enabling the model to better adapt to different image inputs. Text augmentation: Perform operations such as synonym replacement and sentence pattern transformation on the text instructions to increase the diversity and flexibility of the text. These operations can simulate different expression styles and habits of users, enabling the model to better understand natural language instructions;
[0021] Combine with a two-way interaction module for joint enhancement: In the two-way interaction module, the deep interaction between image and text features provides new possibilities for data augmentation. By introducing attention mechanisms and multimodal fusion mechanisms, the model can be more flexible and accurate when processing image and text features. This joint enhancement method can further improve the performance of the model in complex intangible cultural heritage element image editing tasks.
[0022] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0023] 1. The present invention enhances the understanding and reasoning capabilities in instruction editing by combining an MLLM (such as LLaVA). The MLLM can perform collaborative learning across text and image modalities, extracting deep semantic information, enabling the model to not only handle basic instructions but also understand complex intangible cultural heritage elements.
[0024] 2. To improve the model's understanding of intangible cultural heritage elements, we designed an enhanced two-way interaction mechanism. This mechanism realizes deep interaction between image and text features through a cross-attention mechanism, enabling image features to serve as query and key-value pairs for two-way communication with text features.
[0025] 3. To improve the model's performance in complex intangible cultural heritage scenarios, we adopted effective data augmentation strategies. The training data includes not only traditional image editing data but also integrates perception data related to intangible cultural heritage, such as image segmentation data, and synthetic high-quality complex instruction editing data for scenarios requiring world knowledge reasoning. Brief Description of the Drawings
[0026] Figure 1 This is the flowchart of the present invention. Specific implementation manners
[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0028] The present invention provides a technical solution: a modal language model image editing technology integrating intangible cultural heritage (ICH) elements, including:
[0029] Model selection and training: Select the LLaMA model as the basis and introduce LoRA for adaptive fine-tuning. In this way, the model can be adaptively adjusted while keeping the original parameters frozen, greatly improving the training efficiency and instruction understanding ability. Use training data related to ICH for model training.
[0030] Bidirectional interaction module: The bidirectional interaction module includes a feature extraction layer, an interaction layer, and a fusion layer. The feature extraction layer includes text feature extraction and image feature extraction. The text feature extraction uses a pre-trained text encoder to extract the feature vector of the text instruction; the image feature extraction uses a pre-trained image encoder to extract the feature vector of the image. The interaction layer includes text-to-image interaction and image-to-text interaction. The text-to-image interaction inputs the text feature vector as conditional information into the image generation or editing model to affect the image generation or editing process. The image-to-text interaction feeds back the image feature vector to the text processing module to adjust the understanding of the text instruction or generate an image-related text description. The fusion layer fuses the text and image feature vectors output by the interaction layer to generate a fused feature vector for subsequent image generation and editing tasks, realizing multi-round information flow between image and text features. Verify the effectiveness of the bidirectional interaction module through experiments and adjust the model parameters to optimize the performance.
[0031] Diffusion models generate images through step-by-step iteration. Starting from a random noise image, through the step-by-step guidance of instructions and the generation process, the model gradually approaches the target image. Therefore, when dealing with complex intangible cultural heritage image editing tasks, diffusion models can, through their phased generation strategy, ensure that the generated results have sufficient details and precision. Diffusion models can gradually generate these details, making them highly consistent with the original image and conforming to the design style of cultural elements. Combined with the two-way interaction module, diffusion models can not only generate high-quality images that meet user requirements when executing editing instructions, but also ensure the uniqueness of cultural symbols during the image generation process through interaction with text features.
[0032] In the present invention: In step two, the two-way interaction module includes: a feature extraction layer, an interaction layer, and a fusion layer. The feature extraction layer includes text feature extraction and image feature extraction. The text feature extraction uses a pre-trained text encoder to extract the feature vector of the text instruction; the image feature extraction uses a pre-trained image encoder to extract the feature vector of the image; the interaction layer includes text-to-image interaction and image-to-text interaction. The text-to-image interaction inputs the text feature vector as conditional information into the image generation or editing model to affect the image generation or editing process. The image-to-text interaction feeds back the image feature vector to the text processing module to adjust the understanding of the text instruction or generate an image-related text description. The fusion layer fuses the text and image feature vectors output by the interaction layer to generate a fused feature vector for subsequent image generation and editing tasks, realizing multiple rounds of information flow between image and text features. The effectiveness of the two-way interaction module is verified through experiments, and the model parameters are adjusted to optimize the performance.
[0033] The combination of image and text is essentially a process of converting visual information into a language description. In this process, the system first uses a pre-trained image encoder to deeply analyze the image and extract the key feature vectors in the image. These feature vectors not only contain the visual information of the image but also implicitly contain the cultural symbols and connotations carried by the image.
[0034] Subsequently, these image feature vectors are fed back to the text processing module. After receiving these feature vectors, the text processing module will intelligently adjust the text instruction according to the visual information contained in them. This adjustment involves refining, correcting, or supplementing the text instruction to ensure that the text instruction can more accurately reflect the intangible cultural heritage elements and the overall artistic conception in the image.
[0035] Complementary to the combination of image-to-text is the combination of text-to-image. In this process, the system first uses a pre-trained text encoder to parse the text instructions and extract the key semantic information. This semantic information not only contains the user's editing intention but also implies the expectations and requirements for the image content.
[0036] Subsequently, these text feature vectors are input into the image generation or editing model as conditional information. After receiving these feature vectors, the model will generate or edit the image according to the semantic information they contain. This generation or editing not only follows the user's instruction requirements but also incorporates the unique charm and symbols of intangible cultural heritage.
[0037] It realizes the deep interaction between images and text and also brings many advantages. First, this combination significantly improves the intelligence level of image editing, enabling the system to generate or edit images that meet expectations according to the user's instruction requirements. Second, this combination also enhances the semantic expression ability of images, enabling images to more accurately convey the unique charm and symbols of intangible cultural heritage.
[0038] In practical applications, this combination technology can be widely used in the protection and inheritance of intangible cultural heritage, the design and development of the cultural and creative industries, etc. By using this technology, we can easily generate or edit images containing intangible cultural heritage elements, providing strong support for the dissemination and promotion of intangible cultural heritage. At the same time, this technology can also provide new inspiration and ideas for the design and development of the cultural and creative industries, promoting the innovative development of the industry.
[0039] In the present invention: the training data related to intangible cultural heritage in step one includes image segmentation data, complex instruction editing data, and world knowledge reasoning data.
[0040] In the present invention: the two-way interaction module further includes an attention mechanism. The attention mechanism strengthens the association between image and text features. In the text-to-image interaction, the attention mechanism is used to focus on the image regions most relevant to the text instructions; in the image-to-text interaction, the attention mechanism is used to focus on the text parts most relevant to the image features.
[0041] In the present invention: the two-way interaction module further includes a multimodal fusion mechanism. The multimodal fusion mechanism effectively fuses the text and image feature vectors in the fusion layer. The multimodal fusion mechanism includes but is not limited to concatenation, weighted summation, and bilinear pooling. The two-way interaction module can combine the text and image feature vectors together to form a unified multimodal feature representation. The fused multimodal feature representation is used in subsequent image generation and editing tasks. Through this multimodal feature representation, the system can simultaneously understand the text instructions and image content, thereby achieving a more accurate and user-expected image editing result.
[0042] In the present invention, the specific implementation steps of text-to-image interaction are as follows: when a user inputs a natural language instruction containing intangible cultural heritage elements, the system first uses a pre-trained text encoder to extract the feature vector of the text instruction. This feature vector is then input into an image generation or editing model as conditional information, affecting the image generation or editing process. In this way, the descriptions and instructions in the text can directly guide the image generation to ensure that the image content meets the user's expectations. The implementation steps of the image-to-text interaction are as follows: during the image editing process, a pre-trained image encoder is used to extract the feature vector of the image. These image feature vectors are then fed back to the text processing module to adjust the understanding of the text instruction or generate a text description related to the image. This interaction mechanism enables the system to optimize the interpretation of the text instruction according to the actual situation of the image, ensuring that the editing result is more accurate and in line with the user's intention. That is, text feature extraction encodes a natural language instruction using a pre-trained text encoder. The encoder converts the text into a series of high-dimensional feature vectors that can capture the semantic information and context relationships in the text. Image feature extraction encodes an image using a pre-trained image encoder. The encoder converts the image into a series of feature maps or feature vectors that can capture the visual information and structural information in the image.
[0043] In the present invention, the diffusion model includes an initial noise stage, a gradual denoising stage, and a result fusion stage; the result fusion in the result fusion stage is implicitly completed during the iteration of each time step. The model at each time step will perform further denoising and refinement based on the output of the previous time step. Finally, when the preset number of iterations is reached, the model will output a completely denoised and highly refined data sample as the final result. The generation strategy of gradual denoising decomposes complex generation tasks into iterative subtasks at multiple time steps. During the training stage, the model learns how to reverse the forward diffusion process, that is, to recover the original data from the noisy data. During the inference stage, the model randomly samples a noise vector from the standard Gaussian distribution and then gradually removes the noise through the reverse diffusion process to finally generate a clear image or other types of data samples.
[0044] In the present invention, the data augmentation strategies include:
[0045] Introduce multiple data sources related to intangible cultural heritage: Image segmentation data: This type of data contains fine-grained annotations of intangible cultural heritage elements, which helps the model better identify and understand these elements during the learning process. Complex instruction editing data: By providing editing data containing complex instructions, the model can learn how to precisely edit the intangible cultural heritage elements in the image under the guidance of natural language instructions. World knowledge reasoning data: This type of data contains rich cultural background knowledge and common sense information, which helps the model better understand the implicit intentions in the user's instructions during the reasoning process;
[0046] Use diverse data augmentation techniques: Image augmentation: Perform operations such as rotation, scaling, cropping, and color transformation on the image to increase the diversity and robustness of the image. These operations can simulate different shooting conditions and perspectives, enabling the model to better adapt to different image inputs. Text augmentation: Perform operations such as synonym replacement and sentence pattern transformation on the text instructions to increase the diversity and flexibility of the text. These operations can simulate different expression ways and habits of users, enabling the model to better understand natural language instructions;
[0047] Combine with a two-way interaction module for joint augmentation: In the two-way interaction module, the deep interaction between image and text features provides new possibilities for data augmentation. By introducing attention mechanisms and multimodal fusion mechanisms, the model can be more flexible and accurate when processing image and text features. This joint augmentation method can further improve the performance of the model in complex intangible cultural heritage element image editing tasks.
[0048] An image editing technology for a modal language model integrating intangible cultural heritage elements also includes a data augmentation strategy. The data augmentation strategy aims to enrich the model's training data and improve its understanding and processing ability of multimodal inputs. In traditional image editing training data, there usually only include simple image modification or generation tasks, and these data are insufficient when dealing with complex scenarios involving cultural backgrounds. Therefore, we introduce multiple data sources related to intangible cultural heritage for the model's training, including image segmentation data, complex instruction editing data, and world knowledge reasoning data, etc.
[0049] When dealing with multi-object editing, existing technologies often struggle to accurately distinguish target objects, resulting in errors or biases in the editing results. In the "immersive" technology, the model, through its two-way interaction module and the multi-modal collaborative learning ability of MLLM, can accurately identify multiple objects in an image and selectively edit the specified object attributes according to the user's natural language instructions. Compared with existing technologies, the accuracy rate of this model in multi-object editing reaches 92.7%, significantly superior to other baseline methods. When dealing with scenarios involving multiple objects and requiring modification of specific attributes of a certain object (such as color, size, position), traditional methods often confuse the attribute associations between different objects, while "immersive" can ensure that the instructions take effect on the correct target object through its enhanced text-image two-way interaction mechanism. In this way, the model effectively overcomes the limitations of existing technologies in multi-object scenarios and can provide users with more accurate editing effects.
[0050] Since the image editing task of intangible cultural heritage elements often involves more than just simple object editing and requires the model to have a certain reasoning ability, especially in complex scenarios involving cultural background knowledge, the model in the present invention has significantly improved its reasoning ability in complex instruction reasoning tasks. Such tasks require the model to have rich cultural background knowledge and infer the implicit intentions in the user's instructions through cross-modal information interaction. Compared with existing technologies, the reasoning accuracy rate of the model has increased by more than 15%, showing its advantages in complex cultural reasoning tasks. This progress benefits from the combination of MLLM and LoRA fine-tuning, enabling the model to quickly adapt to the requirements of new tasks without affecting the original parameters. In addition, the two-way interaction module also ensures that the model can better reason between image features and text semantics to achieve the accurate execution of complex instructions.
[0051] To verify the application effect of the "immersive" technology in the intangible cultural heritage scenario, we selected the representative intangible cultural heritage element of hairpin flower art and conducted multiple editing task verifications. In these experiments, the model was required to make innovative modifications to the specific details of the hairpin flower according to the instructions. The experimental results show that "immersive" has a high level of cultural understanding and detail processing ability when dealing with intangible cultural heritage scenarios. In the complex and changeable scenarios of intangible cultural heritage symbols, the model can balance the needs of cultural preservation and innovation, and flexibly adjust the image elements through instructions, thus achieving image editing results that conform to the cultural artistic conception and have innovation. This ability not only enhances the application potential of the model in the field of intangible cultural heritage but also opens up a new path for the modern digital inheritance of intangible cultural heritage.
[0052] In summary: The modal language model image editing technology integrating intangible cultural heritage (ICH) elements enhances the understanding and reasoning capabilities in instruction editing by combining MLLMs (such as LLaVA). MLLMs can perform collaborative learning across text and image modalities, extracting deep semantic information, enabling the model to not only handle basic instructions but also understand complex ICH elements. To improve the model's understanding of ICH elements, we designed an enhanced bidirectional interaction mechanism that achieves deep interaction between image and text features through the cross-attention mechanism, allowing image features to serve as query and key-value pairs for two-way communication with text features. To enhance the model's performance in complex ICH scenarios, we adopted an effective data augmentation strategy. The training data includes not only traditional image editing data but also integrates ICH-related perceptual data such as image segmentation data, as well as synthetic high-quality complex instruction editing data for scenarios requiring world knowledge reasoning.
[0053] It should be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.
[0054] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made therein without departing from the principles and spirit of the invention, and the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A modal language model image editing technology integrating intangible cultural elements, characterized in that: The steps include: Step 1: Select the basic model. Select the LLaMA model as the basic model and introduce LoRA for adaptive fine-tuning. While keeping the original parameters of the LLaMA model frozen, introduce LoRA (Low-Rank Adaptation) technology for adaptive fine-tuning. Use the training data related to intangible cultural heritage for model training. Collect text and image data related to intangible cultural heritage for model training. Through training, the model can understand and generate text and images related to intangible cultural heritage. Step 2: construct a two-way interaction module, which includes: a feature extraction layer, an interaction layer and a fusion layer. The feature extraction layer extracts text and image features, the interaction layer realizes the mutual influence between text and image, and the fusion layer generates a fusion feature vector. Step 3: Apply the diffusion model. Starting from a random noise image, the diffusion model generates high-quality images through step-by-step iteration and instruction guidance. The diffusion model adopts a phased strategy to ensure details and accuracy. At the same time, it combines a two-way interactive module to interact with user editing instructions and text features to generate images that meet the requirements and retain the uniqueness of cultural symbols.
2. According to claim 1, a modal language model image editing technology integrating intangible cultural elements is characterized by: The two-way interaction module in step 2 includes: a feature extraction layer, an interaction layer and a fusion layer. The feature extraction layer includes text feature extraction and image feature extraction. The text feature extraction uses a pre-trained text encoder to extract a feature vector of a text instruction; the image feature extraction uses a pre-trained image encoder to extract a feature vector of an image; the interaction layer includes text-to-image interaction and image-to-text interaction. The text-to-image interaction inputs the text feature vector as conditional information into an image generation or editing model to affect the image generation or editing process. The image-to-text interaction feeds back the image feature vector to the text processing module to adjust the understanding of the text instruction or generate a text description related to the image. The fusion layer fuses the text and image feature vectors output by the interaction layer to generate a fused feature vector for subsequent image generation and editing tasks, thereby realizing multiple rounds of information flow between image and text features. The effectiveness of the two-way interaction module is verified through experiments, and the model parameters are adjusted to optimize performance.
3. The modal language model image editing technology integrating intangible cultural elements according to claim 1 is characterized in that: The intangible cultural heritage-related training data in step one includes image segmentation data, complex instruction editing data and world knowledge reasoning data.
4. The modal language model image editing technology integrating intangible cultural elements according to claim 1 is characterized in that: The bidirectional interaction module also includes an attention mechanism, which strengthens the association between image and text features. In text-to-image interaction, the attention mechanism is used to focus on the image area most relevant to the text instruction; in image-to-text interaction, the attention mechanism is used to focus on the text part most relevant to the image feature.
5. The modal language model image editing technology integrating intangible cultural heritage elements according to claim 1 is characterized in that: The bidirectional interaction module also includes a multimodal fusion mechanism, which effectively fuses text and image feature vectors in a fusion layer. The multimodal fusion mechanism includes but is not limited to splicing, weighted summation, and bilinear pooling. The bidirectional interaction module can combine text and image feature vectors together to form a unified multimodal feature representation. The fused multimodal feature representation is used in subsequent image generation and editing tasks. Through this multimodal feature representation, the system can understand text instructions and image content at the same time, thereby achieving more accurate image editing results that meet user expectations.
6. A modal language model image editing technology integrating intangible cultural elements according to any one of claims 1 to 5, characterized in that: It also includes a data enhancement strategy, which introduces multiple data sources related to intangible cultural heritage for model training, including image segmentation data, complex instruction editing data, and world knowledge reasoning data.
7. The modal language model image editing technology integrating intangible cultural elements according to claim 2 is characterized by: The specific implementation steps of the interaction from text to image are: when the user inputs a natural language instruction containing intangible cultural elements, the system first uses a pre-trained text encoder to extract the feature vector of the text instruction, and this feature vector is then input into the image generation or editing model as conditional information to affect the image generation or editing process. In this way, the description and instructions in the text can directly guide the generation of the image to ensure that the image content meets the user's expectations. The image-to-text interaction has the following implementation steps: during the image editing process, the feature vector of the image is extracted using a pre-trained image encoder, and these image feature vectors are then fed back to the text processing module to adjust the understanding of the text instruction or generate a text description related to the image. This interaction mechanism enables the system to optimize the interpretation of the text instruction according to the actual situation of the image, ensuring that the editing result is more accurate and in line with the user's intention, that is, text feature extraction encodes the natural language instruction by using a pre-trained text encoder, and the encoder converts the text into a series of high-dimensional feature vectors, which can capture the semantic information and contextual relationship in the text, and image feature extraction encodes the image by using a pre-trained image encoder, and the encoder converts the image into a series of feature maps or feature vectors, which can capture the visual information and structural information in the image.
8. The modal language model image editing technology integrating intangible cultural elements according to claim 1 is characterized by: The diffusion model is divided into an initial noise stage, a step-by-step denoising stage and a result fusion stage; the fusion in the result fusion stage is implicitly completed during the iteration process of each time step, and the model of each time step will be further denoised and refined based on the output of the previous time step. Finally, when the preset number of iterations is reached, the model will output a completely denoised, highly refined data sample as the final result. The step-by-step denoising generation strategy decomposes the complex generation task into iterative subtasks of multiple time steps. In the training stage, the model learns how to reverse the forward diffusion process, that is, to recover the original data from the noisy data. In the inference stage, the model randomly samples a noise vector from a standard Gaussian distribution, and then gradually removes the noise through the reverse diffusion process, and finally generates a clear image or other types of data samples.
9. The modal language model image editing technology integrating intangible cultural heritage elements according to claim 6 is characterized by: The data enhancement strategy includes: Introduce a variety of data sources related to intangible cultural heritage: Image segmentation data: This type of data contains detailed annotations of intangible cultural heritage elements, which helps the model better identify and understand these elements during the learning process; Complex instruction editing data: By providing editing data containing complex instructions, the model can learn how to accurately edit intangible cultural heritage elements in images under the guidance of natural language instructions; World knowledge reasoning data: This type of data contains rich cultural background knowledge and common sense information, which helps the model better understand the implicit intentions in user instructions during the reasoning process; Use a variety of data enhancement technologies: Image enhancement: Perform operations such as rotation, scaling, cropping, and color transformation on images to increase the diversity and robustness of images. These operations can simulate different shooting conditions and perspectives, allowing the model to better adapt to different image inputs. Text enhancement: Perform operations such as synonym replacement and sentence transformation on text instructions to increase the diversity and flexibility of text. These operations can simulate different expressions and habits of users, allowing the model to better understand natural language instructions. Combined with the bidirectional interaction module for joint enhancement: In the bidirectional interaction module, the deep interaction between image and text features provides new possibilities for data enhancement. By introducing the attention mechanism and multimodal fusion mechanism, the model can be more flexible and accurate in processing image and text features. This joint enhancement method can further improve the performance of the model in complex intangible cultural heritage element image editing tasks.
Citation Information
Cited By
Semantic image editing method and system based on physical perception
CN120876669A
Lake and Hunan woodcarving image generation method, device and equipment based on LoRA model and storage medium
CN121095382A
Lake xiang wood carving image generation method, device and equipment based on LoRA model and storage medium
CN121095382B
Multi-subject personalized image generation method, system and device and storage medium
CN121639859A
Multi-subject personalized image generation method, system, device and storage medium
CN121639859B