Text-guided image refining method and device, equipment and medium
Through the text-guided image refinement method, the problem that the prior art is difficult to take into account overall consistency and local detail accuracy when processing complex image scenes, and achieves higher generation accuracy and efficiency.
Patent Information
- Application Number
- CN202510342404.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-24
Smart Images

Figure CN120198735A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a text-guided image refinement method, device, equipment, and storage medium. Background Art
[0002] Image generation technology has been widely applied in multiple fields of social economy, especially in the fields of medical health and finance. However, when dealing with complex image scenes, existing image generation technologies are difficult to balance overall consistency and the accuracy of local details. Especially in areas involving boundary details and texture features (such as the face, hand, lesion area, etc.), problems such as missing details, distorted structures, and unnatural regional fusions often occur in the generated results. These deficiencies limit the application effects of image generation technology in various practical business scenarios.
[0003] In the field of medical health, image generation technology is widely used for auxiliary diagnosis and medical image analysis. For example, in lesion recognition and organ segmentation tasks, doctors identify boundary positions and key texture features through medical images to accurately judge the lesion situation. However, when generating medical images, existing technologies often face problems such as blurred details in the lesion area, inaccurate boundary recognition, and lack of personalized adaptation for different patient characteristics. These deficiencies may lead to diagnostic errors or missed diagnoses, seriously affecting the accuracy and reliability of medical image analysis.
[0004] In the field of finance, image generation technology has been applied to business scenarios such as customer service and identity authentication, such as generating personalized anthropomorphic images for brand marketing and customer interaction. However, when generating customer anthropomorphic images, existing technologies often have difficulty accurately restoring the facial features and hand details of customers, resulting in distorted generated results and affecting the user experience. At the same time, existing technologies lack the ability to automatically integrate brand features and are difficult to achieve the combination of corporate visual style and customer images. In addition, the financial business has high requirements for the real-time performance of image generation, but existing technologies have high computational costs and low generation efficiency, making it difficult to meet the demand for real-time generation.
[0005] Generally speaking, when dealing with complex image scenes, existing image generation technologies have common problems such as significant deficiencies in regional recognition and refinement, weak adaptive optimization ability, and high computational costs, which limit their wide application in the fields of medical health and finance. How to improve the refinement ability of image generation technology for specific regions, achieve accurate recognition of boundary ranges and texture features, while ensuring generation efficiency and overall image consistency, is a technical problem to be solved currently. Summary of the Invention
[0006] The main objective of the present invention is to provide a text-guided image refinement method, apparatus, device, and storage medium, aiming to solve the technical problem in the prior art that it is difficult to balance the consistency of the overall image and the detail accuracy of specific regions during image generation, resulting in unclear generation effects in local regions.
[0007] To achieve the above objective, the present invention provides a text-guided image refinement method, including:
[0008] Obtain a training data set containing text descriptions and corresponding image data, annotate the structured regions and local significant feature regions in the image data, and generate annotation information for the corresponding structured regions and local significant feature regions;
[0009] Define text trigger keywords, and establish a mapping relationship between each text trigger keyword and the annotation information of the structured regions and local significant feature regions, generating an updated training data set containing the mapping relationship;
[0010] Input the updated training data set into a preliminary generation model, combine the text trigger keywords with the annotation information of the structured regions and local significant feature regions, and identify the structured regions and local significant feature regions corresponding to the text trigger keywords to obtain an identification result;
[0011] Based on the identification result, adjust the attention allocation of the preliminary generation model to the structured regions and local significant feature regions to obtain an optimized generation model;
[0012] Generate a target image containing the refinement result based on the optimized generation model.
[0013] Furthermore, to achieve the above objective, the present invention provides a text-guided image refinement apparatus, including:
[0014] A data annotation module, which obtains a training data set containing text descriptions and corresponding image data, annotates the structured regions and local significant feature regions in the image data, and generates annotation information for the corresponding structured regions and local significant feature regions;
[0015] A keyword mapping module, which defines text trigger keywords and establishes a mapping relationship between each text trigger keyword and the annotation information of the structured regions and local significant feature regions, generating an updated training data set containing the mapping relationship;
[0016] A feature recognition module, which inputs the updated training data set into a preliminary generation model, combines the text trigger keywords with the annotation information of the structured regions and local significant feature regions, and identifies the structured regions and local significant feature regions corresponding to the text trigger keywords to obtain an identification result;
[0017] A model optimization module that, based on the recognition result, adjusts the attention allocation of the preliminary generation model to the structured region and the local significant feature region to obtain an optimized generation model;
[0018] An image generation module that generates a target image containing the refinement result based on the optimized generation model.
[0019] Furthermore, to achieve the above object, the present invention also provides a computer device, which includes a memory, a processor, and a text-guided image refinement program stored in the memory and executable on the processor. When the text-guided image refinement program is executed by the processor, the steps of the text-guided image refinement method as described above are implemented.
[0020] Furthermore, to achieve the above object, the present invention also provides a computer-readable storage medium, on which a text-guided image refinement program is stored. When the text-guided image refinement program is executed by a processor, the steps of the text-guided image refinement method as described above are implemented.
[0021] Beneficial effects: The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as medical and health. It discloses a text-guided image refinement method, including: obtaining a training data set containing text descriptions and image data, annotating specific regions in the image, defining text trigger keywords, and establishing a mapping relationship between the keywords and the annotation information of the specific regions; inputting the updated training data set into a preliminary generation model, identifying the specific regions corresponding to the keywords, and adjusting the attention allocation mechanism of the generation model to optimize the detail generation effect of the model on the specific regions; generating a target image containing the refinement result based on the optimized generation model. By introducing text trigger keywords and combining the annotation information of the significant regions in the training stage, the present invention enhances the model's recognition and optimization capabilities for specific regions, realizes the direct integration of the refinement result into the generator during the generation process, thereby reducing post-processing steps, effectively improving the generation accuracy of specific regions in complex scenarios, ensuring the overall image consistency, while reducing the computational cost and improving the image generation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:
[0023] Figure 1 It is a schematic diagram of an application environment of the text-guided image refinement method in an embodiment of the present invention;
[0024] Figure 2 It is a schematic flowchart of an embodiment of the text-guided image refinement method of the present invention;
[0025] Figure 3 Schematic diagram of functional modules of a preferred embodiment of the text-guided image refinement device of the present invention;
[0026] Figure 4 Schematic diagram of a structure of a computer device in an embodiment of the present invention;
[0027] Figure 5 Another schematic diagram of a structure of a computer device in an embodiment of the present invention. Detailed implementation manners
[0028] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0029] The text-guided image refinement method provided by the embodiments of the present invention can be applied in an application environment such as Figure 1 , where the client communicates with the server through a network. The server can obtain a training data set containing text descriptions and image data through the client, annotate specific regions in the images, define text trigger keywords, and establish a mapping relationship between the keywords and the annotation information of the specific regions; input the updated training data set into a preliminary generation model, identify the specific regions corresponding to the keywords, and adjust the attention allocation mechanism of the generation model to optimize the detail generation effect of the model on the specific regions; generate a target image containing the refinement result based on the optimized generation model. By introducing text trigger keywords and combining the annotation information of significant regions in the training stage, the present invention enhances the model's recognition and optimization capabilities for specific regions, realizes that the refinement result is directly incorporated into the generator during the generation process, thereby reducing post-processing steps, effectively improving the generation accuracy of specific regions in complex scenarios, ensuring the overall image consistency, while reducing the computational cost and improving the image generation efficiency. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smartphones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.
[0030] Please refer to Figure 2 , Figure 2 which is a flowchart of an embodiment of the text-guided image refinement method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described herein may be executed in a different order.
[0031] As Figure 2 shown, the text-guided image refinement method proposed by the present invention includes the following steps:
[0032] S10. Obtain a training data set containing text descriptions and corresponding image data, annotate the structured regions and local significant feature regions in the image data, and generate annotation information for the corresponding structured regions and local significant feature regions.
[0033] In this embodiment, the obtained training data set includes text descriptions and corresponding image data. The text descriptions are used to provide semantic information of the images, and the image data is the input of the generation model. The quality and diversity of the training data directly affect the training effect of the model. Therefore, during the data collection process, it is necessary to obtain representative data samples from multiple sources, including images of different scenarios and different feature types, as well as text descriptions that match the image content. The text descriptions should accurately express the core features of the images, such as the category, location, color, etc. of the objects.
[0034] When implementing, the training data can be obtained through channels such as public data sets, medical image databases, and customer image libraries, and the format of the image data is converted and cleaned to remove blurred, incomplete, or irrelevant image samples. At the same time, natural language processing is performed on the text descriptions to ensure that the description content is clear and the structure is reasonable, facilitating subsequent model training and optimization.
[0035] The purpose of annotation is to identify important regions in the image and attach detailed boundary information and feature information to these regions. Structured regions usually refer to the parts with clear overall contours and boundaries in the image, such as faces, organs, or buildings, etc.; local significant feature regions refer to the local regions in the image that need to be refined, such as textures, edges, or lesion regions. The annotation information of these regions provides a basis for the generation model to refine and optimize.
[0036] When implementing, an approach combining automatic annotation and manual assistance can be adopted. The automatic annotation identifies the boundary and key point features in the image through a saliency detection model and generates preliminary annotation results. For the regions that are not accurately recognized by the model, manual secondary correction can be performed. The annotation results are usually stored in the form of a mask image, where each pixel value represents the category or feature type of the corresponding region, facilitating invocation during the training process.
[0037] The generated annotation information includes the boundary range, position coordinates, and texture features of each region. These information are used to guide the generation model to pay special attention to and optimize these regions during the training process, thereby improving the detail generation effect of the image. The annotation information needs to be associated with the image data so that the model can dynamically adjust the attention allocation according to different regions in the image during training.
[0038] When implementing, the annotation information can be stored as an image mask file and associated with the original image data and text description file. For example, for medical imaging data, the annotation information can be combined with the patient's diagnosis information and imaging data to form a complete training sample; for customer image data in the financial field, the annotation information can be associated with the customer's text description information (such as identity, occupation, interests, etc.) to help the model generate more personalized images.
[0039] Example illustration: In the medical imaging analysis scenario of lesion recognition, doctors analyze the patient's lesion condition through medical imaging data. By obtaining a training data set containing lesion descriptions and imaging data, the boundary positions and texture features of the lesion area are automatically recognized, and detailed annotation information is generated. When training the model, the model can focus on the detailed features of the lesion area, thereby generating a clearer and more accurate lesion image. It can help doctors more accurately identify the lesion area, improve the accuracy and efficiency of diagnosis, and reduce the risk of misdiagnosis or missed diagnosis.
[0040] In the scenario of customer identity authentication or customer portrait generation, financial institutions need to generate anthropomorphic images containing personalized features for customers. By obtaining the customer's description information (such as occupation, interests, gender) and the corresponding image data, the structured areas (such as facial contours) and local significant feature areas (such as eyes, mouth, etc.) in the customer image are automatically annotated. By generating detailed annotation information, the generation model is guided to focus on optimizing the detailed features of these areas, thereby generating a more realistic and accurate customer portrait, enhancing the customer experience, and at the same time improving the recognition accuracy and security of the image during the identity authentication process.
[0041] By obtaining a diverse training data set and accurately annotating the structured areas and local significant feature areas in the image data, the recognition and refinement optimization capabilities of the generation model for important areas in the training stage are improved. It can reduce the post-processing steps and directly optimize the key areas in the training stage, thereby improving the generation accuracy and consistency of the model, reducing the computational cost, and enhancing the application effect of the model in complex image scenarios.
[0042] S20, define text trigger keywords, and establish a mapping relationship between each text trigger keyword and the annotation information of the structured area and the local significant feature area, and generate an updated training data set containing the mapping relationship;
[0043] In this embodiment, the text trigger keyword is an identifier used to guide the image generation model to focus on specific areas in the image. In the image generation task, these keywords are associated with the structured areas and local significant feature areas in the image through the training process, enabling the model to more accurately optimize the detailed features of these areas during the generation process.
[0044] During the implementation process, a set of keywords needs to be defined based on the image content in the training data and the saliency features of specific regions. These keywords can represent different region types, such as "facial region", "hand region", "edge region", etc. The keyword set is usually stored in the form of a list or a database for easy invocation during the training and inference phases. The definition of keywords needs to be combined with the requirements of the actual application scenario and adjusted according to different types of image generation tasks. For example, in the field of medical and health, keywords such as "organ contour" and "lesion area" can be defined; in customer identity recognition in the financial field, keywords such as "facial features" and "hand gesture" can be defined.
[0045] The mapping relationship between text-triggered keywords and image annotation information refers to associating each keyword with the corresponding region information in the image, thereby guiding the model to refine and optimize these regions during the training process. The core of the mapping relationship is to match the embedding vector of the keyword with the annotation information of the image, enabling the model to dynamically adjust the attention allocation mechanism according to the input keyword.
[0046] During the implementation process, first, the image data needs to be annotated to extract the boundary positions and texture features of the structured regions and local salient feature regions. Then, each keyword is converted into a high-dimensional embedding vector and associated with the image annotation information. The storage form of the mapping relationship can be a database in key-value pair format, and the mapping information includes keywords, embedding vectors, and feature descriptions of the annotated regions.
[0047] The updated training dataset is a versioned dataset that contains the mapping relationship between text-triggered keywords and annotation information. By combining the initial training dataset with the mapping relationship, an updated training dataset is generated, enabling the model to identify and process the optimization requirements of specific regions during the training process.
[0048] During the implementation process, first, the initial training dataset is preprocessed to ensure the consistency of the image data and text descriptions. Then, the mapping relationship is integrated into the training dataset, so that each sample contains the corresponding keyword and region annotation information. During the training process, the model can dynamically invoke these mapping relationships to optimize the generation effect of specific region details.
[0049] Example illustration: In medical image analysis, doctors usually need to identify the boundaries and detailed features of lesion areas, such as annotating the boundaries and internal structures of tumors in CT images. By defining keywords such as "lesion area" and "organ boundary", a mapping relationship is established between these keywords and the annotation information in the image data. During the training process, the model can automatically identify the lesion areas in the image and refine the boundary and texture features of this area. By optimizing the generated image, doctors can more accurately judge the lesion situation and improve the diagnostic efficiency and accuracy.
[0050] In financial services, customer identification and personalized services often require generating anthropomorphic images of customers. By defining keywords such as "facial features" and "hand gestures", a mapping relationship is established between these keywords and the annotation information of customer image data. During the image generation process, the model can automatically optimize the facial and hand regions of the customer according to the keywords, making the generated images clearer and more realistic, enhancing the customer experience, and at the same time improving the accuracy and security of image matching during the identification process.
[0051] By defining text trigger keywords and establishing a mapping relationship with the structured regions and local significant feature regions of the image data, the model is refined and optimized for key regions during the training phase. It is possible to directly optimize the generation effect of the model through the training process, reduce post-processing steps, improve generation accuracy and consistency, thereby reducing computational costs and improving image generation efficiency.
[0052] S30, input the updated training dataset into the preliminary generation model, and combine the text trigger keywords with the annotation information of the structured regions and local significant feature regions to identify the structured regions and local significant feature regions corresponding to the text trigger keywords, obtaining an identification result;
[0053] In this embodiment, inputting the updated training dataset into the preliminary generation model enables the model to utilize the mapping relationship between the text trigger keywords and the image annotation information during the training process, thereby identifying the structured regions and local significant feature regions. The updated training dataset includes text descriptions, image data and their annotation information, as well as the mapping relationship between the text trigger keywords and the annotation information, which are the basis for optimizing the generation result during the model training process.
[0054] During implementation, the updated training dataset is loaded into the input end of the model, and the data is input into the generation model in batches for feature extraction. During the input process, it is necessary to ensure that the image data corresponds one-to-one with the text description, and at the same time, the mapping relationship between the keywords and the annotation information can be correctly identified and called in the model. The input layer of the model needs to be able to process both text and image data types and can automatically adjust the attention allocation mechanism according to the mapping relationship.
[0055] Combining the text trigger keywords with the annotation information of the structured regions and local significant feature regions aims to guide the model to accurately locate the key regions in the image during the identification process. The text trigger keywords, as the input guidance information of the model, correspond to the regional features in the annotation information, thereby helping the model identify the boundary positions, key point features, and texture details in the image.
[0056] During the implementation process, the model matches the embedding vectors of text trigger keywords with the regional features in the annotation information. The annotation information includes the boundary range, position coordinates, and texture features of each region, and these features are associated with the keywords through a mapping relationship. During the recognition process, the model dynamically adjusts the attention allocation mechanism, allocating more computational resources to these structured regions and locally significant feature regions to improve the recognition accuracy and the ability to retain details.
[0057] Identifying the structured regions and locally significant feature regions corresponding to text trigger keywords is a key step for the model to generate refined image results. During the recognition process, the model locates and analyzes the structured regions and locally significant feature regions in the image based on the input text trigger keywords and the mapping relationship. The model identifies the boundary positions, key point features, and texture details in the image through the feature extraction layer and the attention allocation mechanism.
[0058] During the implementation process, the model first extracts relevant annotation information based on the embedding vectors of the keywords, and then combines the feature extraction results of the image data to identify the structured regions and locally significant feature regions corresponding to the keywords. The recognition process can be completed by a multi-layer neural network model. For example, a convolutional neural network (CNN) is used for feature extraction, and a self-attention mechanism is used for dynamically allocating attention weights.
[0059] The recognition result is the output result of the model's recognition of the structured regions and locally significant feature regions in the image after inputting and updating the training dataset. The recognition result usually includes the following information: the boundary range and position coordinates of each structured region; the key point features and texture details of each locally significant feature region; the association information between the matched text trigger keywords and the regions.
[0060] The recognition result can be used as an optimization basis in subsequent steps to adjust the generation parameters or the attention allocation mechanism of the model to further improve the image generation effect.
[0061] Example illustration: In medical image analysis, doctors need to identify lesion regions and organ boundaries in images. By inputting text trigger keywords containing lesion descriptions (such as "tumor boundary") and medical image data, the model can automatically identify the lesion boundaries and texture features in the image. The recognition result includes the boundary range and key feature information of each lesion region, helping doctors more accurately judge the morphology and development of the lesions and improving the accuracy and efficiency of diagnosis.
[0062] In customer identity verification and risk assessment in the financial field, customers' facial features and hand features are important verification bases. By inputting text containing customer descriptions to trigger keywords (such as "facial features", "hand gestures") and customer image data, the model can automatically identify the facial boundaries and hand detail features in the customer images. The recognition results include the boundary positions and texture feature information of each feature region, which helps to improve the accuracy of identity verification and the experience of personalized services.
[0063] By inputting the updated training dataset into the preliminary generation model and combining the text-triggered keywords with the annotation information of the structured region and the locally significant feature region for recognition, the recognition accuracy of the model for key regions can be effectively improved.
[0064] S40, based on the recognition results, adjust the attention allocation of the preliminary generation model to the structured region and the locally significant feature region to obtain an optimized generation model;
[0065] In this embodiment, the recognition results are the output data obtained after the model performs feature extraction and matching analysis on the updated training dataset, including the positioning information, boundary range, key point features, and texture details of the structured region and the locally significant feature region. These recognition results are used to guide the model to allocate more computing resources and attention to key regions when generating images, improving the detail performance of these regions.
[0066] In the implementation process, the recognition results are stored in the form of structured data, including the annotation information of each region and the corresponding keyword mapping information. For example, in medical images, the recognition results can include the location, boundary, and lesion texture features of the lesion area; in the financial field, the recognition results can include the boundaries and detail descriptions of customers' facial and hand features.
[0067] Adjusting the attention allocation of the preliminary generation model is to improve the generation accuracy and detail expressiveness of the model in specific regions. The attention allocation mechanism is an important optimization strategy in deep learning models. The model highlights key regions by dynamically adjusting the allocation of computing resources to different regions and reduces the computing of irrelevant regions.
[0068] In the implementation process, the attention allocation is adjusted through the following steps:
[0069] Identify regions with insufficient attention: Compare the recognition results with the significant features in the mapping relationship to find the structured regions and locally significant feature regions where the model has insufficient attention allocation during the initial generation process.
[0070] Calculate the attention weight adjustment value: According to the results of the comparison analysis, determine the attention allocation weight adjustment value for each region. These adjustment values are used to guide the model to allocate more computing resources to key regions in the subsequent generation process.
[0071] Optimize the attention allocation mechanism: Apply the attention weight adjustment value to the attention allocation mechanism of the preliminary generation model to dynamically adjust the model's attention intensity to the structured region and the local significant feature region.
[0072] The optimization process of attention allocation can be achieved by using the multi-head self-attention mechanism or the cross-attention mechanism. These mechanisms can help the model identify important information in the input data and automatically adjust the model's allocation of computing resources to different regions.
[0073] The optimized generation model is a model version with higher detail generation ability after adjusting the attention allocation mechanism. Compared with the preliminary generation model, the optimized generation model can pay more attention to the structured region and the local significant feature region during the generation process, thus generating clearer and more accurate images.
[0074] In the implementation process, perform training iteration operations on the preliminary generation model, use the recognition result and the adjusted attention allocation mechanism as inputs, update the generation parameters of the model, and obtain the optimized generation model. The optimized generation model can be applied to actual scenarios of image generation, such as medical image refinement, customer identity recognition, personalized image generation, etc.
[0075] During the training process, the optimization goal of the model needs to comprehensively consider the overall consistency of the image and the detail accuracy of local features, especially optimize the generation quality of the structured region and the local significant feature region. Each input image generates a corresponding region mask according to the significant region features annotated in the dataset. These region masks are used to weight the loss function, enabling the model to perform more precise refinement optimization on the specified regions during the optimization process.
[0076] The loss function of the model consists of a global loss and a local loss, aiming to balance the overall effect of the image and the optimization of local feature details. The specific total loss function is designed as:
[0077] L = L original + λ1L region1 + λ2L regionl2
[0078] Where: L original represents the global generation loss of the original model, which is used to ensure the overall consistency of the image; L region1 represents the local optimization loss of the structured region; L regionl2 represents the optimization loss of the local significant feature region; λ1 and λ2 are dynamically adjusted weight coefficients, which are weighted according to the importance of each region.
[0079] In practical applications, the model controls the optimization degree of generated details by inputting different text trigger keywords. For example:
[0080] Text description: "Someone is playing the guitar", combined with the text trigger keywords "optimize the facial area" and "optimize the hand area", the model will generate an image with clearer details of the face and hands.
[0081] Text description: "A smiling child", combined with the text trigger keyword "optimize the facial area", the model will pay more attention to the refinement of facial features and the natural presentation of expressions.
[0082] The introduction of keywords enables the model to adjust the regional optimization weights without disrupting the overall fluency of image generation, improve the detail generation effect of specific regions, and ultimately achieve a balance between global consistency and local details.
[0083] Example illustration: In medical image analysis, doctors usually need to identify the boundaries and texture features of lesion areas. By adjusting the attention allocation mechanism of the model, the model can allocate more computing resources to the lesion area, thereby improving the image generation accuracy of the lesion area. The optimized generation model can automatically generate images with clear lesion boundaries and rich texture details, providing more accurate diagnostic basis for doctors.
[0084] In financial services, during customer identity verification and risk assessment, the generation and recognition accuracy of customer images are very important. By optimizing the attention allocation mechanism of the model, the model can pay more attention to the detailed features of the customer's face and hands. When generating customer images, the optimized generation model can refine details such as facial contours and hand postures, improve the accuracy and security of identity verification, and thus enhance the customer experience.
[0085] By adjusting the attention allocation mechanism of the preliminary generation model based on the recognition results, the detail generation ability of the model for structured regions and local significant feature regions can be effectively improved.
[0086] S50. Generate a target image containing the refinement result based on the optimized generation model.
[0087] In this embodiment, the optimized generation model is obtained by adjusting the attention allocation mechanism and generation parameters of the preliminary generation model. This model has higher detail generation ability and can more precisely process structured regions and local significant feature regions in image generation tasks. Compared with the preliminary generation model, the optimized generation model pays more attention to the key regions in the image and can improve the boundary clarity and texture detail performance of these regions.
[0088] In the implementation process, the optimized generation model is usually obtained through multiple rounds of training iterations and attention allocation adjustments. After the model optimization is completed, the model has the ability to refine and generate different regions, and can automatically adjust the processing accuracy of each region according to the text-triggered keywords in the input data.
[0089] Generating the target image containing the refinement results is the core task in the model inference stage. In this process, the optimized generation model refines the structured regions and local significant feature regions according to the input text description and the corresponding image generation target. The refinement results usually include:
[0090] Accurate identification and generation of boundary ranges: enhancing the clarity of the boundaries of structured regions and avoiding blurring or distortion;
[0091] Enhancement of key texture features: refining the texture details of local significant feature regions and avoiding fusion, deformation or missing phenomena.
[0092] In the implementation process, the input text description and text-triggered keywords are input into the optimized generation model. The model automatically refines and optimizes the corresponding regions in the image by identifying the annotation information in the keyword and mapping relationship, and generates the target image containing the refinement results. The output of the target image can be in various formats, such as CT images and X-ray films in medical imaging, or customer identification images in the financial field.
[0093] Example illustration: In the field of medical and health, medical imaging data is an important basis for doctors' diagnosis. However, traditional medical imaging generation models often have problems such as missing details and blurred boundaries when generating lesion regions and organ boundaries, which affect the accuracy of diagnosis. Based on the generation model, it is possible to refine and generate the detailed features of the lesion regions and organ structures, improving the clarity and accuracy of the images.
[0094] The doctor inputs an instruction containing the text description "display the boundary of the lung lesion". The model automatically identifies the lung structure and lesion region in the CT image. Through the optimized generation model, the image generation process can refine the clarity of the lesion boundary, enabling the doctor to more accurately judge the shape, size and development of the lesion, thus improving the diagnostic effect. Especially in the early detection of lung cancer, the model can generate clear images of tumor boundaries to assist doctors in accurate diagnosis.
[0095] In the diagnosis of dermatology, doctors often need to view the detailed features of the lesion area, such as edge changes and skin texture. By inputting the description "refinement of the skin lesion boundary", the model automatically optimizes the boundary and texture details of the lesion area, making the generated image more accurately reflect the lesion situation. This refined image output can assist doctors in more quickly identifying the type and severity of the lesion.
[0096] In customer identity verification and personalized image generation in the financial field, the detailed features of the face and hands are important verification bases. However, when traditional image generation models process customer face features and hand features, there are often problems such as unclear boundaries and missing details, which affect the accuracy of identity verification and the realism of the generated images.
[0097] During the customer identity verification process, by inputting a text description containing "customer face feature refinement", the optimized generation model can automatically identify the face area in the customer image and refine details such as the boundaries of facial features and skin texture. The generated images are clearer and more realistic, which can effectively improve the recognition accuracy of the identity verification system and reduce the risks of misjudgment and identity fraud.
[0098] In personalized services, financial institutions can generate anthropomorphic avatar images for customers. For example, by inputting the text description "smiling customer avatar", the model can automatically optimize details such as the lips and teeth in the customer image, making the generated avatar more natural and realistic. Such personalized avatars can be used in application scenarios such as customer account interfaces and membership cards, enhancing the customer service experience and brand image.
[0099] By generating a target image containing refinement results based on the optimized generation model, the overall effect and local detail performance of image generation can be effectively improved; it can automatically refine the key areas in the image without damaging the overall image smoothness, avoiding the computational cost and unstable effects brought by traditional post-processing methods, thereby improving the image generation efficiency and quality.
[0100] The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as medical and health. It discloses a text-guided image refinement method, including: obtaining a training data set containing text descriptions and image data, annotating specific areas in the images, defining text trigger keywords, and establishing a mapping relationship between the keywords and the annotation information of the specific areas; inputting the updated training data set into a preliminary generation model, identifying the specific areas corresponding to the keywords, and adjusting the attention allocation mechanism of the generation model to optimize the detail generation effect of the model on the specific areas; generating a target image containing refinement results based on the optimized generation model. By introducing text trigger keywords and combining the annotation information of significant areas in the training stage, the present invention enhances the model's recognition and optimization capabilities for specific areas, realizes the direct integration of refinement results into the generator during the generation process, thereby reducing post-processing steps and effectively improving the generation accuracy of specific areas in complex scenarios.
[0101] In one embodiment, the above S10 includes:
[0102] S101. Obtain an initial training dataset containing text descriptions and corresponding image data;
[0103] S102. Identify the boundary positions and key point features of the structured regions and local significant feature regions in the image data through a saliency detection model;
[0104] S103. Generate saliency region information based on the identified boundary positions and key point features, where the saliency region information is used to identify the boundary ranges and key texture features of the structured regions and local significant feature regions;
[0105] S104. Generate annotation information for the corresponding structured regions and local significant feature regions based on the saliency region information.
[0106] In this embodiment, the training dataset contains a large amount of text descriptions and corresponding image data, which are used to guide the model to identify and annotate key regions in the image during the training phase. The text descriptions are used to express the content information of the image, while the image data provides visual features to facilitate the model to learn the feature distribution and detail performance of different regions in the image. During the acquisition process, it is necessary to ensure the matching degree between the text description and the image data so that the model can accurately understand the meaning and features of different regions.
[0107] When identifying the structured regions and local significant feature regions in the image data through a saliency detection model, the model needs to focus on analyzing the changes in the boundary positions and key point features in the image. The boundary positions are used to describe the contour shapes of the regions in the image, while the key point features are used to identify the texture changes and detail distributions inside the regions. These information play a crucial role in the identification and generation processes of the model, and can significantly improve the generation accuracy of the detailed regions of the image.
[0108] After identifying the boundary positions and key point features, it is necessary to generate saliency region information. The saliency region information includes the specific descriptions of the boundary ranges of the regions and the key texture features. These information are used to identify the structured regions and local significant feature regions, and guide the model to pay more attention to the detail performance of these regions during the model generation process, avoiding generating blurred or distorted image content.
[0109] Based on the saliency region information, generate annotation information for the corresponding structured regions and local significant feature regions. The annotation information contains detailed descriptions such as the category, position, size, and texture features of the regions, which are called by the model during the training and generation processes. The accuracy of the annotation information directly affects the detail performance effect of the model during the image generation process. Therefore, when generating the annotation information, multiple rounds of verification and adjustment are required to ensure that it can accurately reflect the actual feature distribution in the image.
[0110] In this embodiment, by obtaining the training datasets of text descriptions and image data and annotating the structured regions and local salient feature regions, the recognition ability of the model for key regions and the refinement generation ability can be effectively improved. Compared with the traditional manual annotation method, the automated annotation process can significantly improve the annotation efficiency and accuracy, enabling the model to better learn the key region features in the image during the training process, so as to automatically optimize the detail performance of these regions during the generation process, reduce post-processing operations, and improve the consistency and detail clarity of image generation.
[0111] In one embodiment, before the above S102, it further includes:
[0112] S1011, performing image denoising processing on the image data to remove the noise and artifacts in the image data;
[0113] S1012, performing color space conversion on the image data to convert the image data into a grayscale image for feature recognition;
[0114] S1013, performing edge enhancement processing on the image data and inputting the image data after edge enhancement processing into the saliency detection model.
[0115] In this embodiment, the image denoising processing aims to remove the random noise and artifacts in the image data, and these interferences may come from factors such as image acquisition devices and environmental light changes. If the denoising processing is not performed, the model may be affected by noise when recognizing the structured regions and local salient feature regions, resulting in inaccurate recognition of the boundary positions and key point features. Therefore, the denoising processing is one of the preprocessing steps for the input image of the saliency detection model.
[0116] When implementing the denoising processing, various denoising algorithms can be used, such as:
[0117] Gaussian filtering, which is used to smooth the image and reduce random noise;
[0118] Median filtering, which is used to remove salt-and-pepper noise and is especially suitable for images containing strong bright or dark points;
[0119] Bilateral filtering, which can keep the clarity of the image edges while removing noise and is suitable for scenarios with high requirements for boundary recognition.
[0120] Color space conversion is the process of converting the image data from a color image to a grayscale image. The grayscale image only retains the brightness information of the image while removing the color information, which can simplify the computational complexity of the model and focus on the recognition of the structured features and local salient features of the image.
[0121] The specific implementation of color space conversion includes:
[0122] RGB to grayscale conversion, which converts the red, green, and blue channels of an image into a single grayscale value;
[0123] Brightness weighting method, which performs weighted conversion according to the human eye's perception sensitivity to different colors, retaining important edge and texture information in the image.
[0124] The converted grayscale image can more clearly display the boundary and detail features in the image, facilitating feature extraction by the saliency detection model.
[0125] Edge enhancement processing aims to highlight the edge features in the image, enabling the model to more accurately identify the boundary positions and key point features in the image. Edges are significant changes between different regions in the image and usually contain important structural information. If the edges are not clear, the model may have difficulty accurately identifying the region boundaries, resulting in unsatisfactory refinement generation results.
[0126] The implementation of edge enhancement processing can adopt the following technical means:
[0127] Sobel operator, which calculates the image gradient to identify the edge positions in the image;
[0128] Laplacian operator, which is used to highlight the regions with drastic brightness changes in the image, further enhancing the edge features;
[0129] Canny edge detection algorithm, which combines Gaussian filtering and edge tracing to generate a clear edge image.
[0130] The image data after edge enhancement processing contains clearer boundary information, which can effectively improve the recognition accuracy of the saliency detection model for structured regions and locally significant feature regions.
[0131] In this embodiment, through image denoising processing, color space conversion, and edge enhancement processing, the input image quality of the saliency detection model can be effectively improved, enabling the model to more accurately identify the structured regions and locally significant feature regions in the image. It can not only reduce noise interference but also highlight the boundary information of the image, enabling the model to better restore the image details during the recognition and generation processes. This preprocessing process can significantly improve the effect of image refinement generation, making the generated images more accurate and efficient in applications in fields such as healthcare and finance.
[0132] In one embodiment, the above S20 includes:
[0133] S201, defining a set of text trigger keywords, where each text trigger keyword in the set of text trigger keywords corresponds to a target region, and the target region includes a structured region or a locally significant feature region;
[0134] S202. Encode each text trigger keyword into an embedding vector, which is used to represent the association information between the corresponding text trigger keyword and the target region;
[0135] S203. Extract the salient features of the target region based on the annotation information of the structured region and the local salient feature region;
[0136] S204. Associate the embedding vector with the salient features of the target region to establish a mapping relationship between each text trigger keyword and the annotation information;
[0137] S205. Update the initial training dataset based on the mapping relationship to generate an updated training dataset containing the mapping relationship.
[0138] In this embodiment, the text trigger keyword set is a group of keywords used to guide the generation model to focus on specific image regions. These keywords appear in natural language form, such as "facial details", "hand gestures", etc., corresponding to the target regions that the model needs to optimize during the generation process. The target regions include structured regions (such as lesions in medical images, field regions in tabular data) and local salient feature regions (such as facial features and hand edges in portrait images).
[0139] Each text trigger keyword corresponds to a target region, and this correspondence helps the generation model to selectively optimize the image detail regions in specific tasks, improving the accuracy of the generation results. For example, the text keyword "eye details" corresponds to the eye region in the human face, and the text keyword "lesion boundary" corresponds to the lesion region in the medical image.
[0140] First, define the keyword set and ensure the diversity and pertinence of the keywords. For different application scenarios, specific keyword sets can be defined. For example:
[0141] In the medical field, keywords such as "lung nodule boundary", "organ contour" can be defined;
[0142] In the financial field, keywords such as "facial details of ID card photo", "handwritten signature area" can be defined.
[0143] Then, establish the correspondence between each keyword and the target region to ensure that the model can correctly identify and optimize the relevant regions during the generation process according to the keywords.
[0144] The text trigger keywords are input into the generation model in the form of embedding vectors. Embedding vectors convert natural language keywords into numerical representations that can be understood by computers. This numerical representation contains the semantic information of the keywords and the association information of the target regions. In actual operation, pre-trained natural language processing models (such as BERT, Word2Vec) can be used to encode the keywords.
[0145] Through the embedding vectors, the model can automatically understand the semantic meanings of the keywords and the corresponding target regions, and thus allocate more attention to the image regions related to the keywords during the generation process.
[0146] The generation process of the keyword embedding vectors includes the following steps:
[0147] Use a pre-trained natural language processing model to convert the keywords into vector representations;
[0148] Align the vector representation with the annotation information of the target region to ensure that the vector contains the significant feature information of the target region;
[0149] Input the generated embedding vectors into the generation model to guide the attention allocation of the model during training and generation.
[0150] The significant features of the target region are the region features that the model needs to pay special attention to during training. These features usually include detailed information such as boundary positions, key points, and texture features, which are used to guide the model to optimize the detailed performance of image generation. By extracting the significant features, the model can more accurately identify the key regions in the image and avoid blurring or distortion during the generation process.
[0151] The extraction of significant features needs to be combined with the annotation information to ensure the accuracy and consistency of the feature data. For example, when identifying the facial region, the significant features may include the boundaries and texture features of regions such as eyes, mouth, and nose; when identifying medical images, the significant features may include features such as the shape and density of the edges of lesions.
[0152] The extraction of significant features includes the following steps:
[0153] Use a saliency detection model to extract features from the image data;
[0154] Combine the annotation information to identify the boundary positions and key texture features of the target region;
[0155] Integrate the extracted significant feature information with the annotation data of the target region to ensure that the model can call this feature information during training and generation.
[0156] The correlation between the embedded vectors and the salient features is used to guide the model's attention to different regions during the training and generation processes. By establishing a mapping relationship, the model can automatically identify relevant target regions based on the input text trigger keywords and allocate more computing resources to these regions, thereby enhancing the detailed performance of image generation.
[0157] The process of establishing the mapping relationship ensures the semantic correspondence between the keywords and the target regions, enabling the model to adjust the attention allocation to the image regions according to different keywords. For example, the keyword "eye details" will guide the model to allocate more attention to the eye region, while the keyword "signature area" will guide the model to focus on the details of the handwritten signature.
[0158] The process of establishing the mapping relationship includes the following steps:
[0159] Compare the embedded vectors of the text trigger keywords with the salient features of the target regions;
[0160] Based on the comparison results, establish the correspondence between the keywords and the salient features;
[0161] Store the mapping relationship in the training dataset for the model to call during the training and generation processes.
[0162] The updated training dataset is the dataset after integrating the initial training dataset and the keyword mapping relationship, which contains all the information required by the model during the training process. This information includes image data, text descriptions, annotation information, and keyword mapping relationships, which can help the model better learn the correspondence between text and images during the training process.
[0163] The generation of the updated training dataset ensures that the model can automatically optimize the detailed regions of the image according to the input text description, thereby improving the quality and accuracy of image generation.
[0164] The process of generating the updated training dataset includes the following steps:
[0165] Integrate the text descriptions and image data in the initial training dataset with the keyword mapping relationship;
[0166] Verify whether the integrated dataset contains complete annotation information and mapping relationships;
[0167] Input the updated dataset into the model for training.
[0168] For example, the text trigger keywords are bound to specific local region data through the training process. The specific process is as follows:
[0169] Each text trigger keyword corresponds to a unique embedded vector. For example, the embedded vector of "<focus - structured region>" is represented as Zstructural and the embedding vector representation of "<Focus - Significant Feature Region>" is Z salient .
[0170] During the text encoding process, the text trigger keywords are processed by an encoder (such as the CLIP model) into high - dimensional embedding vectors, which are combined with the embedding vectors describing the global image content to form the following formula:
[0171] Z combined = Z text + αZ structural + βZ salient
[0172] where α and β are coefficients that control the weights of the text trigger keywords and are used to determine the degree of attention of the model to the structured region and the significant feature region.
[0173] During the training process, by collecting diverse high - resolution data and making detailed annotations for key regions. The structured region annotations are used to cover the regular parts in the image, such as tables, forms, or other well - organized regions; the significant feature region annotations are used to mark specific details in the image, such as key features like edges and textures. The mask binds the text trigger keywords to specific regions, enabling the model to allocate more computational resources to these key regions through the attention mechanism, thereby improving the generation effect.
[0174] In this embodiment, through the similarity analysis based on the visual embedding and the label text embedding vector, the system can effectively measure the semantic correlation degree between the image data and the label text, providing basic support for the subsequent heatmap generation; it can capture the deep - level semantic correlation between the image and the text, improving the semantic alignment ability in multi - modal tasks and the accuracy of text - guided image refinement.
[0175] In one embodiment, the above S30 includes:
[0176] S301, input the updated training dataset into the preliminary generation model;
[0177] S302, perform feature extraction operations on the image data in the updated training dataset through the preliminary generation model to extract the basic features of the structured region and the local significant feature region in the image data;
[0178] S303, based on the embedding vector of the text trigger keyword, perform matching analysis on the basic features, identify the structured region and the local significant feature region that match the text trigger keyword, and generate an identification result;
[0179] S304. Based on the significant features of the structured regions and local significant feature regions included in the mapping relationship, perform localization and refinement processing on the boundary range and key texture features of the recognition result to obtain the refined recognition result.
[0180] In this embodiment, the updated training dataset is a dataset after associating text trigger keywords with image annotation information, including text descriptions, image data, annotation information of structured regions and local significant feature regions, and keyword mapping relationships. Inputting this dataset into the preliminary generation model is to utilize the feature extraction ability of the model to extract the basic features of structured regions and local significant feature regions from the image data, providing basic data for subsequent matching analysis and refinement processing.
[0181] Load the updated training dataset into the input end of the model to ensure that the format of the dataset matches the input format of the model; initialize the parameters of the model to ensure that the feature extraction module can operate normally; call the forward propagation process of the model to start the feature extraction operation.
[0182] The feature extraction operation is one of the core steps of the generation model. By performing feature extraction on the image data through the preliminary generation model, the basic structure and local features in the image can be identified, providing support for subsequent keyword matching analysis. Feature extraction includes operations such as boundary recognition, key point localization, and texture analysis to ensure that the basic features of structured regions and local significant feature regions can be accurately captured.
[0183] Extract the boundary information and texture features of the image through a convolutional neural network (CNN); use a saliency detection algorithm to identify information such as the framework of structured regions and table lines; perform key point detection on local significant feature regions (such as facial features and hand details) to extract the boundaries and texture features of these regions.
[0184] The matching analysis is to compare the embedding vector of the text trigger keyword with the basic features of the image to identify the structured regions and local significant feature regions in the image related to the keyword. This process ensures that the model can automatically locate and identify the target regions in the image according to the input text description, thereby improving the generation ability and refinement effect of the model.
[0185] Encode the text trigger keyword into an embedding vector through a pre-trained text encoder (such as the CLIP model); calculate the cosine similarity between the embedding vector and the image feature vector to evaluate the correlation between the two; identify the structured regions and local significant feature regions that match the text trigger keyword according to the similarity score; store the matching result as the recognition result for subsequent refinement processing.
[0186] The saliency features in the mapping relationship include the saliency features of the structured region and the locally significant feature region, such as the boundary range, key texture points, etc. These feature information are used to guide the model to further refine the recognition result, ensuring that the image generation result is more accurate and realistic in terms of detail performance. The refinement process includes operations such as boundary optimization, texture enhancement, noise removal, etc., making the recognition result closer to the actual image details.
[0187] Read the boundary range and key texture features included in the mapping relationship; optimize the boundary position of the recognition result to remove unnecessary edge noise; enhance the key texture features of the recognition result to ensure the clarity and realism of local details; generate the refined recognition result and store it as the output result of the model.
[0188] In this embodiment, by inputting the updated training dataset into the preliminary generation model and combining the text trigger keywords with the annotation information of the structured region and the locally significant feature region for recognition and refinement processing, the accuracy and fineness of the generation model in the image refinement process can be significantly improved; it can automatically identify the keyword information in the text description and allocate the attention of the model to the target region related to the keyword, thereby realizing the automatic optimization and accurate generation of the image detail region.
[0189] In one embodiment, the above S40 includes:
[0190] S401, compare and analyze the recognition result with the saliency features of the structured region and the locally significant feature region in the mapping relationship to generate the identification information of the region with insufficient attention allocation;
[0191] S402, based on the identification information, determine the attention weight adjustment value of the structured region and the locally significant feature region of each region with insufficient attention allocation;
[0192] S403, optimize the attention allocation mechanism of the preliminary generation model through the attention weight adjustment value to adjust the processing weights of the preliminary generation model in the structured region and the locally significant feature region;
[0193] S404, based on the recognition result and the optimized attention allocation mechanism, perform training iteration operations on the preliminary generation model to adjust the generation parameters of the preliminary generation model to obtain the optimized generation model.
[0194] In this embodiment, the purpose of comparing and analyzing the recognition result with the saliency features in the mapping relationship is to find out the detail regions that the generation model fails to fully focus on in the structured region and the locally significant feature region. The comparison and analysis identify the regions with insufficient attention allocation by calculating the difference between the recognition result and the saliency features. These insufficient regions may include unclear boundaries, missing texture details, etc.
[0195] Read the boundary ranges and texture features of the structured regions and locally significant feature regions included in the recognition results; extract the corresponding saliency features from the mapping relationship, including boundary positions and key texture features; compare the recognition results with the saliency features in the mapping relationship, calculate the differences between them, and identify the regions with insufficient attention allocation; mark the difference regions as regions with insufficient attention allocation and generate corresponding identification information for subsequent adjustment.
[0196] The attention weight adjustment value is used to guide the generation model to allocate more computing resources to the regions with insufficient attention allocation during the training process. The adjustment value determined according to the identification information can effectively improve the processing accuracy of the model for boundary ranges and key texture features, and avoid problems such as missing details or distortion in the generated results.
[0197] According to the identification information of the regions with insufficient attention allocation, identify the types and features of each insufficient region; set different attention weight adjustment strategies for different types of regions (such as structured regions or locally significant feature regions); calculate the weight adjustment value for each region, and the weight adjustment value can be dynamically adjusted according to the size of the difference; output the adjustment value for each region with insufficient attention allocation to guide the optimization of the model's attention allocation.
[0198] The attention allocation mechanism is an important part of the generation model, which determines the proportion of computing resources allocated by the model to different regions during the image generation process. By optimizing the attention allocation mechanism, it can be ensured that the model is more accurate in processing structured regions and locally significant feature regions, allocate more computing resources to key regions, and thus improve the detail performance of the generated images.
[0199] Read the attention allocation mechanism parameters of the preliminary generation model; update the model's attention allocation mechanism according to the attention weight adjustment value to make the model allocate higher attention to structured regions and locally significant feature regions; ensure that the optimized attention allocation mechanism can adjust the processing weights of the model in different regions in real time and improve the detail accuracy of the generated images; verify the applicability of the optimized attention allocation mechanism to different input data.
[0200] The training iteration operation is a training process through multiple loops to continuously optimize the generation parameters of the model to ensure that the model can accurately generate the target image containing refined features. After each iterative training, the generation parameters of the model will be updated according to the new data and adjustment values, so that the generation model is more accurate in processing structured regions and locally significant feature regions.
[0201] Based on the recognition results and the optimized attention allocation mechanism, reconstruct the training dataset; perform training iteration operations on the preliminary generation model, and optimize the generation parameters of the model according to the new weight adjustment values during each iteration; monitor the change of the loss value during the training process to ensure the continuous improvement of the generation effect of the model; after the training iteration ends, output the optimized generation model, which can generate target images containing refined features more accurately.
[0202] In this embodiment, by adjusting the attention allocation mechanism of the preliminary generation model based on the recognition results, the ability of the generation model to refine the structured region and the local significant feature region can be significantly improved. By dynamically adjusting the attention weights of the model and optimizing the generation parameters of the model, the model can more accurately identify and optimize the key regions in the image, avoiding the problem of missing details in the generation results.
[0203] In one embodiment, the above S50 includes:
[0204] S501, receive input data including text description and text trigger keywords;
[0205] S502, input the input data into the optimized generation model;
[0206] S503, based on the optimized generation model, generate a target image with refined features in the structured region and the local significant feature region.
[0207] In this embodiment, receiving input data including text description and text trigger keywords is the starting step of the image generation process. The text description provides the semantic information of the image content that the user hopes to generate, while the text trigger keywords are used to guide the generation model to refine specific regions. The core of this process is to convert the user's input data into an encoded form that can be understood and processed by the model.
[0208] Extract the text description and text trigger keywords from the user input; use the natural language processing (NLP) module to encode the text description into a text embedding vector; extract and verify the text trigger keywords to ensure that the format and content of the keywords meet the requirements of the generation model; use the text description and text trigger keywords as a set of input data and prepare to input them into the optimized generation model.
[0209] Inputting the text description and text trigger keywords into the optimized generation model is one of the core steps of image generation. The optimized generation model contains generation parameters and an attention allocation mechanism adjusted through training iterations, and can automatically generate images containing refined features according to the input data.
[0210] Input the embedding vectors of the text description and the text trigger keywords into the encoder part of the generation model; call the attention allocation mechanism of the model to identify the key regions in the input data according to the text trigger keywords; combine the generation parameters of the model and the optimized attention allocation mechanism to complete the forward propagation process of the model; transfer the output vector of the model to the decoder part to generate image data containing refined features.
[0211] The core of generating the target image is to ensure that the structured regions and local significant feature regions in the image have higher detail expressiveness. The model dynamically adjusts the attention allocation to the key regions during the image generation process by combining the guidance of the text description and the text trigger keywords, so as to generate the target image containing refined features.
[0212] In the decoder part of the model, convert the output vector into image data; according to the attention allocation mechanism of the model, further refine the structured regions and local significant feature regions in the image data; perform refinement operations such as boundary optimization and texture enhancement to make the details of the target image in the structured regions and local significant feature regions clearer and more natural; output the final target image and save the image in the format required by the user, such as JPEG, PNG, etc.
[0213] In this embodiment, by generating the target image containing the refinement result based on the optimized generation model, the detail expressiveness and clarity of the generated image in the key regions can be significantly improved. By dynamically adjusting the attention allocation mechanism of the model, the automatic optimization of the refined features of the image is realized, reducing the need for post - repair or manual adjustment.
[0214] In one embodiment, a text - guided image refinement device is provided, and the text - guided image refinement device corresponds one - to - one with the text - guided image refinement method in the above - mentioned embodiment. Refer to Figure 3 , Figure 3 is a schematic diagram of the functional modules of a preferred embodiment of the text - guided image refinement device of the present invention. Data annotation module 10, keyword mapping module 20, feature recognition module 30, model optimization module 40, and image generation module 50. The detailed description of each functional module is as follows:
[0215] The data annotation module 10 obtains a training data set containing text descriptions and corresponding image data, annotates the structured regions and local significant feature regions in the image data, and generates annotation information corresponding to the structured regions and local significant feature regions.
[0216] The keyword mapping module 20 defines text trigger keywords and establishes a mapping relationship between each text trigger keyword and the annotation information of the structured regions and local significant feature regions, generating an updated training data set containing the mapping relationship.
[0217] The feature recognition module 30 inputs the updated training data set into the preliminary generation model, combines the text trigger keywords with the annotation information of the structured region and the local significant feature region, and recognizes the structured region and the local significant feature region corresponding to the text trigger keywords to obtain a recognition result;
[0218] The model optimization module 40 adjusts the attention allocation of the preliminary generation model to the structured region and the local significant feature region based on the recognition result to obtain an optimized generation model;
[0219] The image generation module 50 generates a target image containing the refinement result based on the optimized generation model.
[0220] In one embodiment, the data annotation module 10 is specifically configured to:
[0221] Obtain an initial training data set containing text descriptions and corresponding image data;
[0222] Identify the boundary positions and key point features of the structured region and the local significant feature region in the image data through a saliency detection model;
[0223] Generate saliency region information according to the identified boundary positions and key point features, where the saliency region information is used to identify the boundary range and key texture features of the structured region and the local significant feature region;
[0224] Generate corresponding annotation information for the structured region and the local significant feature region based on the saliency region information.
[0225] In one embodiment, the data annotation module 10 is specifically configured to:
[0226] Perform image denoising processing on the image data to remove noise and artifacts in the image data;
[0227] Perform color space conversion on the image data to convert the image data into a grayscale image for feature recognition;
[0228] Perform edge enhancement processing on the image data and input the image data after edge enhancement processing into the saliency detection model.
[0229] In one embodiment, the keyword mapping module 20 is specifically configured to:
[0230] Define a set of text trigger keywords, where each text trigger keyword in the set of text trigger keywords corresponds to a target region, and the target region includes a structured region or a local significant feature region;
[0231] Encode each text trigger keyword into an embedding vector, where the embedding vector is used to represent the association information between the corresponding text trigger keyword and the target region;
[0232] Extract the saliency features of the target region based on the annotation information of the structured region and the local salient feature region;
[0233] Associate the embedding vector with the saliency features of the target region to establish a mapping relationship between each text trigger keyword and the annotation information;
[0234] Update the initial training dataset based on the mapping relationship to generate an updated training dataset containing the mapping relationship.
[0235] In one embodiment, the feature recognition module 30 is specifically configured to:
[0236] Input the updated training dataset into the preliminary generation model;
[0237] Perform feature extraction operations on the image data in the updated training dataset through the preliminary generation model to extract the basic features of the structured region and the local salient feature region in the image data;
[0238] Based on the embedding vector of the text trigger keyword, perform matching analysis on the basic features, identify the structured region and the local salient feature region that match the text trigger keyword, and generate an identification result;
[0239] Through the saliency features of the structured region and the local salient feature region included in the mapping relationship, perform positioning and refinement processing on the boundary range and key texture features of the identification result to obtain the refined identification result.
[0240] In one embodiment, the model optimization module 40 is specifically configured to:
[0241] Compare and analyze the identification result with the saliency features of the structured region and the local salient feature region in the mapping relationship to generate identification information for the region with insufficient attention allocation;
[0242] Based on the identification information, determine the attention weight adjustment value of the structured region and the local salient feature region of each region with insufficient attention allocation;
[0243] Optimize the attention allocation mechanism of the preliminary generation model through the attention weight adjustment value to adjust the processing weights of the preliminary generation model in the structured region and the local salient feature region;
[0244] Based on the identification result and the optimized attention allocation mechanism, perform training iteration operations on the preliminary generation model to adjust the generation parameters of the preliminary generation model to obtain an optimized generation model.
[0245] In one embodiment, the image generation module 50 is specifically configured to:
[0246] Receive input data including a text description and a text trigger keyword;
[0247] Input the input data into the optimized generation model;
[0248] Based on the optimized generation model, generate a target image with refined features in the structured region and the local significant feature region.
[0249] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as Figure 4 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a text-guided image refinement method.
[0250] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as Figure 5 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a text-guided image refinement method
[0251] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:
[0252] Obtain a training data set including a text description and corresponding image data, label the structured region and the local significant feature region in the image data, and generate corresponding annotation information for the structured region and the local significant feature region;
[0253] Define text trigger keywords, and establish a mapping relationship between each text trigger keyword and the annotation information of the structured region and the local significant feature region, and generate an updated training dataset containing the mapping relationship;
[0254] Input the updated training dataset into the preliminary generation model, and combine the text trigger keywords with the annotation information of the structured region and the local significant feature region to identify the structured region and the local significant feature region corresponding to the text trigger keywords, and obtain the recognition result;
[0255] Based on the recognition result, adjust the attention allocation of the preliminary generation model to the structured region and the local significant feature region to obtain an optimized generation model;
[0256] Generate a target image containing the refinement result based on the optimized generation model.
[0257] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0258] Obtain a training dataset containing text descriptions and corresponding image data, annotate the structured region and the local significant feature region in the image data, and generate the annotation information of the corresponding structured region and the local significant feature region;
[0259] Define text trigger keywords, and establish a mapping relationship between each text trigger keyword and the annotation information of the structured region and the local significant feature region, and generate an updated training dataset containing the mapping relationship;
[0260] Input the updated training dataset into the preliminary generation model, and combine the text trigger keywords with the annotation information of the structured region and the local significant feature region to identify the structured region and the local significant feature region corresponding to the text trigger keywords, and obtain the recognition result;
[0261] Based on the recognition result, adjust the attention allocation of the preliminary generation model to the structured region and the local significant feature region to obtain an optimized generation model;
[0262] Generate a target image containing the refinement result based on the optimized generation model.
[0263] It should be noted that for the functions or steps that can be realized by the above computer-readable storage medium or computer device, reference can be made to the relevant descriptions on the server side and the user side in the foregoing method embodiments. To avoid repetition, they will not be described in detail here.
[0264] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0265] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0266] It should be noted that if non-company software tools or components appear in the embodiments of the present application, they are only used for illustrative introduction and do not represent actual use. The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A text-guided image refinement method, characterized in that: The following steps are involved: Acquire a training data set including text descriptions and corresponding image data, annotate structured regions and local significant feature regions in the image data, and generate corresponding annotation information of the structured regions and the local significant feature regions; Defining text trigger keywords, and establishing a mapping relationship between each text trigger keyword and the annotation information of the structured region and the local significant feature region, and generating an updated training data set containing the mapping relationship; Inputting the updated training data set into the preliminary generation model, combining the text trigger keywords with the annotation information of the structured regions and the local significant feature regions, identifying the structured regions and the local significant feature regions corresponding to the text trigger keywords, and obtaining the recognition results; Based on the recognition result, adjusting the attention allocation of the preliminary generation model to the structured area and the local significant feature area to obtain an optimized generation model; A target image including a refinement result is generated based on the optimized generation model.
2. The text-guided image thinning method according to claim 1, characterized in that: Acquiring a training data set including text descriptions and corresponding image data, annotating structured regions and local significant feature regions in the image data, and generating corresponding annotation information of the structured regions and local significant feature regions, including: Obtain an initial training dataset containing text descriptions and corresponding image data; Identifying the boundary positions and key point features of the structured regions and local salient feature regions in the image data by using a saliency detection model; Generate salient region information according to the identified boundary position and key point features, wherein the salient region information is used to identify the boundary range and key texture features of the structured region and the local salient feature region; Based on the salient region information, labeling information of corresponding structured regions and local salient feature regions is generated.
3. The text-guided image thinning method according to claim 2, characterized in that: Before identifying the boundary positions and key point features of the structured regions and local salient feature regions in the image data by using a saliency detection model, the method further includes: Performing image denoising processing on the image data to remove noise and artifacts in the image data; Performing color space conversion on the image data to convert the image data into a grayscale image for feature recognition; The image data is subjected to edge enhancement processing, and the image data after edge enhancement processing is input into the saliency detection model.
4. The text-guided image thinning method according to claim 1, characterized in that: Defining text trigger keywords, and establishing a mapping relationship between each text trigger keyword and the annotation information of the structured region and the local significant feature region, and generating an updated training data set containing the mapping relationship, including: Define a text trigger keyword set, each text trigger keyword in the text trigger keyword set corresponds to a target area, and the target area includes a structured area or a local significant feature area; Encoding each text trigger keyword into an embedding vector, wherein the embedding vector is used to represent association information between the corresponding text trigger keyword and the target area; Extract the salient features of the target area based on the annotation information of the structured area and the local salient feature area; Associating the embedding vector with the salient features of the target area to establish a mapping relationship between each text trigger keyword and the annotation information; The initial training data set is updated based on the mapping relationship to generate an updated training data set including the mapping relationship.
5. The text-guided image thinning method according to claim 1, characterized in that: The updated training data set is input into the preliminary generation model, and the text trigger keywords are combined with the annotation information of the structured region and the local significant feature region to identify the structured region and the local significant feature region corresponding to the text trigger keywords, and obtain the recognition results, including: Input the updated training data set into the preliminary generation model; Performing a feature extraction operation on the image data in the updated training data set through the preliminary generation model to extract basic features of structured regions and local significant feature regions in the image data; Based on the embedding vector of the text trigger keyword, matching analysis is performed on the basic features to identify the structured regions and local significant feature regions that match the text trigger keyword, and generate recognition results; The boundary range and key texture features of the recognition result are positioned and refined by using the salient features of the structured region and the local salient feature region contained in the mapping relationship to obtain a recognition result after refinement.
6. The text-guided image thinning method according to claim 1, characterized in that: Based on the recognition result, the attention allocation of the preliminary generation model to the structured area and the local significant feature area is adjusted to obtain an optimized generation model, including: Compare and analyze the recognition result with the salient features of the structured region and the local salient feature region in the mapping relationship to generate identification information of the insufficient attention allocation region; Based on the identification information, determining the attention weight adjustment value of the structured area and the local salient feature area of each insufficient attention allocation area; Optimizing the attention allocation mechanism of the preliminary generation model by using the attention weight adjustment value to adjust the processing weights of the preliminary generation model in the structured area and the local significant feature area; Based on the recognition results and the optimized attention allocation mechanism, a training iteration operation is performed on the preliminary generation model to adjust the generation parameters of the preliminary generation model to obtain an optimized generation model.
7. The text-guided image thinning method according to claim 1, characterized in that: Generating a target image including a refinement result based on the optimized generation model includes: receiving input data including a text description and a text trigger keyword; Inputting the input data into the optimized generation model; Based on the optimized generation model, a target image with refined features in structured areas and local salient feature areas is generated.
8. A text-guided image refinement device, characterized in that: The text-guided image refinement device comprises: A data annotation module obtains a training data set including text descriptions and corresponding image data, annotates structured regions and local significant feature regions in the image data, and generates annotation information of corresponding structured regions and local significant feature regions; A keyword mapping module defines text trigger keywords, establishes a mapping relationship between each text trigger keyword and the annotation information of the structured region and the local significant feature region, and generates an updated training data set containing the mapping relationship; The feature recognition module inputs the updated training data set into the preliminary generation model, combines the text trigger keywords with the annotation information of the structured regions and the local significant feature regions, recognizes the structured regions and the local significant feature regions corresponding to the text trigger keywords, and obtains the recognition results; A model optimization module, based on the recognition result, adjusts the attention allocation of the preliminary generation model to the structured area and the local significant feature area to obtain an optimized generation model; An image generation module generates a target image including a refinement result based on the optimized generation model.
9. A computer device, characterized in that: The computer device comprises a memory, a processor, and a text-guided image refinement program stored in the memory and executable on the processor, wherein the text-guided image refinement program implements the steps of the text-guided image refinement method according to any one of claims 1 to 7 when executed by the processor.
10. A computer-readable storage medium, characterized in that: The storage medium stores a text-guided image thinning program, and when the text-guided image thinning program is executed by the processor, the steps of the text-guided image thinning method according to any one of claims 1 to 7 are implemented.