Method and device for improving integrating degree of image service, equipment and medium
By parsing the semantic relationships in the input text and constructing semantic enhancement constraints, the problem of object and attribute misalignment in image generation in the Diffusion model is solved, achieving higher image generation accuracy and business fit, suitable for financial and medical scenarios.
Patent Information
- Application Number
- CN202510707589.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-12
AI Technical Summary
The Diffusion model may cause misalignment of objects and attributes during the image generation process, resulting in images that do not meet expectations and affecting the effectiveness of business applications in financial and medical scenarios.
By parsing the semantic relationship between each object and each attribute in the input text, a hierarchical binding relationship is generated, semantic enhancement constraints are constructed, and the binding relationship is maintained through the loss function during the image generation process. The trained Diffusion model is used for image generation.
It improves the accuracy and business relevance of image generation, enhances work efficiency, reduces retraining costs, and provides more reliable technical support for financial and medical scenarios.
Smart Images

Figure CN120635307A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology for financial and medical scenarios, and in particular to a method, apparatus, device, and medium for improving the business compatibility of images. Background Art
[0002] In the field of image processing technology for financial and medical scenarios, the precise generation of image materials is crucial for business compatibility. Diffusion models are often used to generate image materials, but these models can cause misalignment of objects and attributes during the image generation process. For example, in financial scenarios, when images with specific logos and item combinations need to be generated, traditional methods can cause misalignment, resulting in images that do not conform to expectations and impacting business applications. Similarly, in medical scenarios, accurate image generation is crucial for critical processes such as diagnosis and treatment, and misalignment can also have significant impacts. Summary of the Invention
[0003] The present invention provides an artificial intelligence method, device, computer equipment and medium for improving the business compatibility of images, so as to solve the problem of object and attribute misalignment that may occur when a diffusion model generates images.
[0004] In a first aspect, a method for improving the business compatibility of an image is provided, comprising:
[0005] Parse the semantic relationship between each object and each attribute in the input text, and generate a binding relationship between objects and corresponding attributes with a hierarchical structure;
[0006] Constructing semantic enhancement constraints based on the binding relationship;
[0007] The input text and semantic enhancement constraints are input into the trained Diffusion model for image generation processing, and the binding relationship in the input text is maintained through a loss function during the image generation process to output a target image with spatial and semantic matching.
[0008] In a second aspect, a device for improving image service compatibility is provided, comprising:
[0009] The relationship binding module is used to analyze the semantic relationship between each object and each attribute in the input text and generate a binding relationship between the object and the corresponding attribute with a hierarchical structure;
[0010] A constraint building module, used for building semantic enhancement constraint conditions based on the binding relationship;
[0011] The image processing module is used to input the input text and semantic enhancement constraints into the trained Diffusion model for image generation processing, and maintain the binding relationship in the input text through a loss function during the image generation process to output a target image with spatial and semantic matching.
[0012] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned method for improving image business compatibility when executing the computer program.
[0013] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned method for improving image business compatibility are implemented.
[0014] In the solution implemented by the above-mentioned method, device, computer equipment and storage medium for improving the business compatibility of images, a binding relationship between objects and corresponding attributes with a hierarchical structure can be generated by parsing the semantic relationship between each object and each attribute in the input text; semantic enhancement constraints are constructed based on the binding relationship; the input text and semantic enhancement constraints are input into the trained Diffusion model for image generation processing, and the binding relationship in the input text is maintained through a loss function during the image generation process to output a target image with spatial and semantic matching. In the present invention, by optimizing the training strategy, the problem of misalignment of objects and attributes that may occur when the Diffusion model generates images is effectively solved, thereby improving the accuracy of image generation and business compatibility. This improvement not only improves work efficiency, but also significantly reduces the cost caused by retraining, providing more reliable technical support for image processing in financial and medical scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0016] Figure 1 This is a schematic diagram of an application environment of a method for improving image service compatibility in one embodiment of the present invention;
[0017] Figure 2 This is a flow chart of a method for improving image service compatibility in one embodiment of the present invention;
[0018] Figure 3 yes Figure 2A schematic flow chart of a specific implementation of step S101;
[0019] Figure 4 yes Figure 2 A schematic flow chart of a specific implementation of step S102;
[0020] Figure 5 yes Figure 2 A schematic flow chart of a specific implementation of the loss function in step S103;
[0021] Figure 6 This is a flowchart of a specific implementation of attention graph alignment in one embodiment of the present invention;
[0022] Figure 7 1 is a schematic structural diagram of an apparatus for improving image service compatibility in one embodiment of the present invention;
[0023] Figure 8 is a structural diagram of a computer device in one embodiment of the present invention;
[0024] Figure 9 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0026] The method for improving image service compatibility provided by the embodiment of the present invention can be applied in the following aspects: Figure 1In an application environment, a client communicates with a server via a network. The server can receive input text from the client and parse the semantic relationships between objects and attributes in the input text to generate hierarchical binding relationships between objects and corresponding attributes; construct semantic enhancement constraints based on the binding relationships; input the input text and semantic enhancement constraints into a trained Diffusion model for image generation processing; and maintain the binding relationships in the input text through a loss function during the image generation process to output a target image that matches both spatially and semantically. In the present invention, by optimizing the training strategy, the problem of object and attribute misalignment that may occur when the Diffusion model generates images is effectively resolved, thereby improving the accuracy and business relevance of image generation. This improvement not only improves work efficiency but also significantly reduces the cost of retraining, providing more reliable technical support for image processing in financial and medical scenarios. The client can include, but is not limited to, various personal computers, laptops, smartphones, tablet computers, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.
[0027] See also Figure 2 As shown, Figure 2 A flowchart of a method for improving image service compatibility provided by an embodiment of the present invention includes the following steps S101-S103.
[0028] S101, parsing the semantic relationship between each object and each attribute in the input text, generating a binding relationship between the object and the corresponding attribute in a hierarchical structure;
[0029] In step S101, precise semantic analysis accurately captures the inherent connections between objects and attributes in the input text, ensuring the accuracy of the generated binding relationships. This hierarchical binding relationship not only helps maintain semantic consistency during subsequent image generation, but also effectively avoids image generation errors caused by misalignment of objects and attributes, significantly improving image generation accuracy and business relevance.
[0030] In one embodiment, if Figure 3 As shown, step S101 includes:
[0031] S201, extracting a text component tree of the input text through dependency syntax analysis and component syntax analysis, and determining the central word of each object and its associated attribute word set;
[0032] S202: Establishing attribute binding priorities based on dependency distances between object center words and attribute words, and generating binding relationships between objects and corresponding attributes with a hierarchical structure.
[0033] In steps S201-S202, by combining dependency parsing and component parsing, we can gain a deeper understanding of the input text structure, ensuring that the core word and its related attribute words for each object are accurately identified. This detailed analysis helps to establish more accurate and stable object-attribute binding relationships in subsequent steps.
[0034] For example, in the financial field, a user may need to generate a promotional image for a financial product. The input text describes various attributes of the financial product (object), such as yield, risk, etc. Through dependency parsing and component parsing, the central word "financial product" can be accurately identified, and attribute words associated with it, such as "yield", "risk", etc., can be found. Then, based on the dependency distance between these attribute words and the central word of the financial product, the priority of attribute binding can be established. For example, "yield" may be regarded as a more important attribute, so it can be placed in a more prominent position when generating the image. In this way, the generated image can not only accurately reflect the characteristics of the financial product, but also better meet the needs of users and business scenarios, thereby improving the business fit of the image.
[0035] S102. Constructing semantic enhancement constraints based on binding relationships;
[0036] In step S102, by introducing semantic enhancement constraints, the image generation process can be further optimized to ensure that the image content is highly consistent with user intent and business scenarios. These constraints can be set based on factors such as the priority of the object-attribute binding relationship, the range of attribute values, or specific requirements. For example, for a promotional image of a financial product, it can be set that the yield must be highlighted and its numerical range must be consistent with the actual product; at the same time, the risk level can be intuitively represented by different colors or icons so that users can quickly understand it. Through such detailed semantic enhancement constraints, the generated image will be more in line with user expectations and business needs, thereby effectively improving the business fit of the image and user experience.
[0037] In one embodiment, if Figure 4 As shown, step S102 includes:
[0038] S301, generating a binary position mask for each object in the image space, marking a preset distribution area of the object;
[0039] S302: Construct an attribute association matrix to record the binding strength value between each attribute word and the corresponding object, wherein the strength value is determined according to the grammatical analysis result.
[0040] In steps S301-S302, by generating a binary position mask for each object in the image space, the preset distribution area of the object in the image can be accurately marked, which provides clear object location information for subsequent image generation. At the same time, an attribute association matrix is constructed to record the binding strength value of each attribute word and the corresponding object, which can quantify the degree of association between the attribute word and the object, ensuring that the attribute word can be accurately and reasonably assigned to the corresponding object when the image is generated. Such technical processing not only improves the accuracy and efficiency of image generation, but also better ensures the consistency of image content with user intent and business scenarios, thereby further improving the business fit of the image and user experience.
[0041] For example, in a financial field, when a user needs to generate a promotional image for a financial product, the positions of key elements such as product name, yield, and risk level in the image can be pre-set through step S301, such as placing the product name in the top center of the image, and the yield and risk level are prominently distributed below the product name. Subsequently, in step S302, by constructing an attribute association matrix, it can be ensured that attribute words such as "high yield" and "low risk" can be accurately bound to the corresponding product image elements, such as strongly binding "high yield" to a graphic element that displays a specific numerical value, and strongly binding "low risk" to an icon element that represents safety or stability. This processing method ensures that the generated image not only meets the promotional standards of financial products, but also can intuitively and accurately convey the characteristics of the product, thereby effectively attracting user attention and enhancing user favorability and trust in the product.
[0042] For example, in the medical field, when a user needs to generate an image of a medical report or disease introduction, step S301 can be used to pre-plan the layout of key information in the image, such as the disease name, symptom description, and treatment recommendations. For example, the disease name is placed in the center above the image, the symptom description is arranged in a list below the disease name, and the treatment recommendations are displayed next to the symptom description in the form of an icon or flowchart. Then, in step S302, an attribute association matrix is constructed to ensure that attribute words such as "serious symptoms" and "effective treatment" are closely associated with the corresponding image elements. For example, "serious symptoms" are strongly bound to graphics or color codes that describe the severity of the symptoms, and "effective treatment" is strongly bound to icons that display specific treatment plans or drugs. This processing method not only makes the generated image structure clear and the information hierarchy distinct, but also ensures the accuracy and authority of medical information, helping patients better understand the disease condition. It also provides medical staff with an intuitive and efficient communication tool, thereby improving the overall quality of medical services and patient satisfaction.
[0043] S103, inputting the input text and the semantic enhancement constraint condition into the trained Diffusion model for image generation processing, and maintaining the binding relationship in the input text through the loss function during the image generation process to output a target image with spatial and semantic matching;
[0044] In step S103, by combining the input text and semantic enhancement constraints, the Diffusion model can more accurately capture the key information in the text during the image generation process and create images based on this information. The application of the loss function ensures that the binding relationship in the text is maintained during the image generation process, that is, the objects, attributes and their mutual relationships described in the text are accurately presented in the generated image. This process not only improves the fit between the image and the text content, but also makes the generated image more visually in line with the user's expectations, enhancing the readability and communication efficiency of the image information. At the same time, due to the introduction of semantic enhancement constraints, the generated image is also richer and more accurate at the semantic level, providing users with a more comprehensive and in-depth image information experience.
[0045] In one embodiment, if Figure 5 As shown, the step of maintaining the binding relationship in the input text by using the loss function during the image generation process in step S103 includes:
[0046] S401, maximizing the similarity between the object and the corresponding attribute through a loss function;
[0047] S402, minimizing the similarity between the object and the non-corresponding attribute through a loss function;
[0048] S403: Minimize the similarity between the attribute and the non-corresponding object through a loss function.
[0049] In steps S401-S403, a refined loss function design ensures a close match between objects and their corresponding attributes during image generation, while effectively avoiding erroneous associations between objects and non-corresponding attributes, and vice versa. This design not only improves image generation accuracy but also results in finer detail in the generated images, with clearer and more explicit relationships between objects and attributes.
[0050] For example, in a financial field, if a user wants to generate an image describing a certain financial product, the input text may contain key information such as the name, type, and income characteristics of the product. Through step S401, the loss function will ensure that the similarity between the object in the image (such as the financial product icon) and its corresponding attribute (such as high return, low risk) is maximized, so that the image can intuitively reflect the core characteristics of the product. In step S402, the loss function will minimize the similarity between the object and the non-corresponding attribute to avoid the appearance of attribute descriptions in the image that are inconsistent with the product characteristics, such as mistakenly associating high-risk attributes with a financial product known for its stability. In step S403, the loss function will further ensure that the similarity between the attribute and the non-corresponding object is minimized to prevent the attribute from being mistakenly applied to other unrelated objects, thereby maintaining the accuracy of the binding relationship described in the text. This refined loss function design is particularly important in scenarios such as the financial field where highly accurate information communication is required.
[0051] Specifically, the loss function L can be constructed based on attention map alignment, and the specific formula is:
[0052] L=α·Σ(S_o,a-max(S_o,a')+β·Σ(S_a,o-max(S_a,o');
[0053] Among them, S_o,a represents the similarity of the attention map between object o and the corresponding attribute a, S_o,a' represents the similarity between object o and non-corresponding attribute a', S_a,o' represents the similarity between attribute a and non-corresponding object o', and α and β are weight coefficients.
[0054] More specifically, Figure 6 As shown, the steps of attention map alignment may include:
[0055] S501, extracting the cross-layer feature map corresponding to the object center word in the decoder of the Diffusion model;
[0056] S502, calculating the cosine similarity matrix between the corresponding attribute words and the cross-layer feature map;
[0057] S503: Constrain the similarity matrix based on the attribute association matrix, and strengthen the feature associations with binding strength values higher than a threshold.
[0058] The method based on steps S501-S503 can more accurately locate the position of the object center word in the image and the association between it and the attribute words. By calculating the cosine similarity matrix, the similarity between the object and the attribute can be quantified, providing a basis for subsequent attribute alignment. Constraining the similarity matrix based on the attribute association matrix can further filter out strongly bound feature associations, weaken or exclude those interference items with weaker binding strength, thereby ensuring that the final generated image description can accurately and efficiently convey the core attributes of the object and improve the fit between the image and the business.
[0059] For example, in the medical field, if a medical image description is needed, this method can be used to precisely locate the lesion region (as the object center word) and further associate it with attribute words such as "size," "shape," and "location." By extracting the cross-layer feature maps corresponding to the lesion region in the Diffusion model decoder and calculating the cosine similarity matrix between these feature maps and the attribute words, the similarity between the lesion and the attributes can be quantified. Next, based on a pre-constructed attribute association matrix, the similarity matrix is constrained to emphasize attribute features closely related to the lesion, such as the lesion's "irregular shape" or "specific location," while weakening or excluding attributes with weaker associations, such as "uniform color." This results in a medical image description that more accurately conveys the lesion's core attributes, helping doctors make quick and accurate diagnoses and improving the alignment between medical imaging and clinical practice.
[0060] In one embodiment, the method for improving image business compatibility also includes: monitoring the binding status of the object and the corresponding attribute, and when it is detected that the spatial offset exceeds a preset allowable threshold, re-injecting the initial constraint parameters and resetting the attenuation progress of the attention adjustment mechanism to the initial state.
[0061] This embodiment can ensure that the binding relationship between objects and corresponding attributes always remains within a reasonable range, avoiding inaccurate attribute descriptions caused by spatial offsets. Through real-time monitoring and adjustment, this method can dynamically maintain the association between objects and attributes in the image, ensuring that the generated image description is highly consistent with the actual situation, thereby further improving the image business fit. This mechanism is particularly important in complex or dynamically changing image scenes, and can significantly improve the accuracy and reliability of image descriptions.
[0062] As can be seen from the above solution, the present invention utilizes a method based on Linguistic Binding to improve the business relevance of images. By optimizing the training strategy, the present invention effectively addresses the object and attribute misalignment issue that can occur when generating images using the Diffusion model, thereby improving image generation accuracy and business relevance. This improvement not only increases work efficiency but also significantly reduces retraining costs, providing more reliable technical support for image processing in financial and medical scenarios.
[0063] Furthermore, because the present invention performs semantic loss calculations, this technology eliminates the need for additional image datasets. Instead, it requires pre-compilation of a sufficiently large text string dataset using tools. With the assistance of other tools, sentence structure parsing and annotation are then performed. This is then trained in a Diffusion model to produce a model. When the model is implemented in a business context, the same parsing tools and annotation rules are used for parsing and object and attribute location annotation. After this processing, the input text is fed into the trained model to produce an image that meets the business requirements.
[0064] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0065] In one embodiment, a device for improving image service compatibility is provided, and the device for improving image service compatibility corresponds to the method for improving image service compatibility in the above embodiment. Figure 6 As shown, the device for improving image business compatibility includes a relationship binding module 601, a constraint construction module 602 and an image processing module 603. The functional modules are described in detail as follows:
[0066] The relationship binding module 601 is used to analyze the semantic relationship between each object and each attribute in the input text, and generate a binding relationship between the object and the corresponding attribute with a hierarchical structure;
[0067] A constraint construction module 602 is used to construct semantic enhancement constraint conditions based on binding relationships;
[0068] The image processing module 603 is used to input the input text and semantic enhancement constraints into the trained Diffusion model for image generation processing, and maintain the binding relationship in the input text through the loss function during the image generation process to output a target image with spatial and semantic matching.
[0069] In one embodiment, the relationship binding module 601 is specifically configured to:
[0070] Extract the text component tree of the input text through dependency syntax analysis and component syntax analysis, and determine the central word of each object and its associated attribute word set;
[0071] The attribute binding priority is established based on the dependency distance between the object center word and the attribute word, and the binding relationship between the object and the corresponding attribute with a hierarchical structure is generated.
[0072] In one embodiment, the constraint construction module 602 is specifically configured to:
[0073] Generate a binary position mask for each object in the image space, marking the preset distribution area of the object;
[0074] Construct an attribute association matrix to record the binding strength value between each attribute word and the corresponding object, where the strength value is determined according to the grammatical analysis result.
[0075] In one embodiment, the image processing module 603 maintains the binding relationship in the input text through a loss function during the image generation process, specifically for:
[0076] Maximize the similarity between the object and the corresponding attribute through the loss function;
[0077] Minimize the similarity between objects and non-corresponding attributes through the loss function;
[0078] The loss function is used to minimize the similarity between attributes and non-corresponding objects.
[0079] Among them, the loss function L is constructed based on the attention map alignment, and the specific formula is:
[0080] L=α·Σ(S_o,a-max(S_o,a'))+β·Σ(S_a,o-max(S_a,o'));
[0081] Among them, S_o,a represents the similarity of the attention map between object o and the corresponding attribute a, S_o,a' represents the similarity between object o and non-corresponding attribute a', S_a,o' represents the similarity between attribute a and non-corresponding object o', and α and β are weight coefficients.
[0082] In one embodiment, the process of attention map alignment specifically includes:
[0083] Extract the cross-layer feature map corresponding to the object center word in the decoder of the Diffusion model;
[0084] Calculate the cosine similarity matrix between the corresponding attribute words and the cross-layer feature maps;
[0085] The similarity matrix is constrained based on the attribute association matrix to strengthen the feature associations with binding strength values above the threshold.
[0086] In one embodiment, the apparatus for improving image service compatibility is further configured to:
[0087] Monitor the binding status of objects and corresponding attributes. When it is detected that the spatial offset exceeds the preset allowable threshold, the initial constraint parameters are re-injected and the attenuation progress of the attention adjustment mechanism is reset to the initial state.
[0088] The device for improving image business suitability, provided by this invention, effectively addresses the object and attribute misalignment issue that can occur when generating images using the Diffusion model by optimizing training strategies, thereby improving image generation accuracy and business suitability. This improvement not only increases work efficiency but also significantly reduces retraining costs, providing more reliable technical support for image processing in financial and medical scenarios.
[0089] For the specific definition of the device for improving the image service compatibility, please refer to the definition of the method for improving the image service compatibility above, which will not be repeated here. The various modules in the above-mentioned device for improving the image service compatibility can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0090] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 8 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the service side of a method for improving the compatibility of image services.
[0091] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 9As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the client side of a method for improving the compatibility of image services.
[0092] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0093] Parse the semantic relationship between each object and each attribute in the input text, and generate a binding relationship between objects and corresponding attributes with a hierarchical structure;
[0094] Construct semantically enhanced constraints based on binding relationships;
[0095] The input text and semantic enhancement constraints are input into the trained Diffusion model for image generation processing. During the image generation process, the binding relationship in the input text is maintained through the loss function to output the target image with spatial and semantic matching.
[0096] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0097] Parse the semantic relationship between each object and each attribute in the input text, and generate a binding relationship between objects and corresponding attributes with a hierarchical structure;
[0098] Construct semantically enhanced constraints based on binding relationships;
[0099] The input text and semantic enhancement constraints are input into the trained Diffusion model for image generation processing. During the image generation process, the binding relationship in the input text is maintained through the loss function to output the target image with spatial and semantic matching.
[0100] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0101] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM) and memory bus dynamic RAM (RDRAM).
[0102] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0103] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A method for improving image business compatibility, characterized in that: include: Parse the semantic relationship between each object and each attribute in the input text, and generate a binding relationship between objects and corresponding attributes with a hierarchical structure; Constructing semantic enhancement constraints based on the binding relationship; The input text and semantic enhancement constraints are input into the trained Diffusion model for image generation processing, and the binding relationship in the input text is maintained through a loss function during the image generation process to output a target image with spatial and semantic matching.
2. The method for improving image business compatibility according to claim 1, wherein: The step of parsing the semantic relationship between each object and each attribute in the input text to generate a binding relationship between the object and the corresponding attribute in a hierarchical structure includes: Extracting a text component tree of the input text through dependency syntax analysis and component syntax analysis to determine a central word of each object and a set of associated attribute words; The attribute binding priority is established based on the dependency distance between the object center word and the attribute word, and the binding relationship between the object and the corresponding attribute with a hierarchical structure is generated.
3. The method for improving image business compatibility according to claim 1, wherein: The step of constructing a semantic enhancement constraint condition based on the binding relationship includes: Generate a binary position mask for each object in the image space, marking the preset distribution area of the object; An attribute association matrix is constructed to record the binding strength value between each attribute word and the corresponding object, wherein the strength value is determined according to the grammatical analysis result.
4. The method for improving image business compatibility according to claim 3, wherein: Maintaining the binding relationship in the input text by using a loss function during the image generation process includes: Maximize the similarity between the object and the corresponding attribute through the loss function; Minimize the similarity between objects and non-corresponding attributes through the loss function; The loss function is used to minimize the similarity between attributes and non-corresponding objects.
5. The method for improving image business compatibility according to claim 4, wherein: The loss function L is constructed based on attention map alignment, and the specific formula is: L=α·Σ(S_o,a-max(S_o,a')+β·Σ(S_a,o-max(S_a,o'); Among them, S_o,a represents the similarity of the attention map between object o and the corresponding attribute a, S_o,a' represents the similarity between object o and non-corresponding attribute a', S_a,o' represents the similarity between attribute a and non-corresponding object o', and α and β are weight coefficients.
6. The method for improving image business compatibility according to claim 5, wherein: The step of aligning the attention map includes: Extracting the cross-layer feature map corresponding to the object center word in the decoder of the Diffusion model; Calculating a cosine similarity matrix between corresponding attribute words and the cross-layer feature map; The similarity matrix is constrained based on the attribute association matrix, and feature associations with binding strength values higher than a threshold are strengthened.
7. The method for improving image business compatibility according to claim 1, wherein: Also includes: Monitor the binding status of objects and corresponding attributes. When it is detected that the spatial offset exceeds the preset allowable threshold, the initial constraint parameters are re-injected and the attenuation progress of the attention adjustment mechanism is reset to the initial state.
8. A device for improving image business compatibility, characterized in that: include: The relationship binding module is used to analyze the semantic relationship between each object and each attribute in the input text and generate a binding relationship between the object and the corresponding attribute with a hierarchical structure; A constraint building module, used for building semantic enhancement constraint conditions based on the binding relationship; The image processing module is used to input the input text and semantic enhancement constraints into the trained Diffusion model for image generation processing, and maintain the binding relationship in the input text through a loss function during the image generation process to output a target image with spatial and semantic matching.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method for improving image business compatibility as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for improving image business compatibility as claimed in any one of claims 1 to 7 are implemented.