Quality evaluation method and device for character embedded in image and storage medium

By using the text aesthetic evaluation model and positional relationship analysis in the image embedded text evaluation, the problem of inaccurate text evaluation is solved, and a more accurate and comprehensive quality evaluation is achieved.

CN120259209APending Publication Date: 2025-07-04CHINA UNITED NETWORK COMM GRP CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510315132.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing text evaluation methods cannot accurately evaluate the quality of text embedded in the image, resulting in inaccurate evaluation results.

Method used

By acquiring the target image, the pre-trained text aesthetic evaluation model is used to evaluate the aesthetics of embedded text, combining the positional relationship and text layout between embedded text and image entities, and weighted summing processing is used to obtain the quality evaluation results of embedded text.

Benefits of technology

It realizes a comprehensive and accurate evaluation of the quality of embedded text in the image, comprehensively considering aesthetics, positional relationships and layout, and improves the accuracy and comprehensiveness of the evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259209A_ABST
    Figure CN120259209A_ABST
Patent Text Reader

Abstract

The invention provides an image embedded character quality evaluation method and device and a storage medium. The method comprises the steps that a target image is acquired, and the target image is an image containing embedded characters; based on the target image, according to a pre-trained character aesthetics evaluation model, obtaining an aesthetics score of the embedded character; obtaining a character embedding position score of the embedded character according to a position relationship between the embedded character and an entity in the target image; determining a character layout score of the embedded character based on the target image; and performing weighted summation processing on the attractiveness score, the character embedding position score and the character layout score of the embedded character to obtain a quality evaluation result of the embedded character. According to the method provided by the invention, the quality of the embedded characters in the image can be accurately evaluated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a method, device, and storage medium for evaluating the quality of embedded text in images. Background Art

[0002] With the development of large model technology, text-to-image large models have made great progress in image quality and generation speed, and have been substantially applied in many fields such as advertising, clothing, animation, and news. In addition to generating simple pictures, there is a demand in fields such as posters, advertisements, ink paintings, and product designs to embed text in the picture, such as poems in ink paintings, themes in posters, slogans in advertisements, and text printed on clothes.

[0003] However, there is currently no quality evaluation scheme for embedded text in images. If existing text evaluation methods are directly used for evaluating embedded text in images, there is a problem that the evaluation results are inaccurate. Summary of the Invention

[0004] This application provides a method, device, and storage medium for evaluating the quality of embedded text in images to solve the technical problem of inaccurate evaluation of embedded text in images.

[0005] In a first aspect, this application provides a method for evaluating the quality of embedded text in images, including:

[0006] Obtain a target image, where the target image is an image containing embedded text;

[0007] Based on the target image, obtain the aesthetics score of the embedded text according to a pre-trained text aesthetics evaluation model;

[0008] Obtain the text embedding position score of the embedded text according to the positional relationship between the embedded text and the entities in the target image;

[0009] Based on the target image, determine the text layout score of the embedded text;

[0010] Perform weighted summation processing on the aesthetics score, text embedding position score, and text layout score of the embedded text to obtain the quality evaluation result of the embedded text.

[0011] Optionally, in the above method, obtaining the text embedding position score of the embedded text according to the positional relationship between the embedded text and the entities in the target image includes:

[0012] According to the target image, determine the first bounding box of the embedded text and the second bounding box of the non-text entities in the target image;

[0013] Obtain the text embedding position score of the embedded text according to the positional relationship between the first bounding box and the second bounding box.

[0014] Optionally, for the method as above, obtaining the text embedding position score of the embedded text according to the positional relationship between the first bounding box and the second bounding box includes:

[0015] If the first bounding box is within the range of the second bounding box, then according to the entity category corresponding to the second bounding box containing the first bounding box, obtain the reasonableness of the embedded text being on the entity corresponding to the entity category by interacting with the large language model, and determine the text embedding position score of the embedded text according to the reasonableness.

[0016] If the first bounding box is outside the range of the second bounding box, then obtain the reasonableness of the embedded text being on a non-entity by interacting with the large language model, and determine the text embedding position score of the embedded text according to the reasonableness.

[0017] If there is an intersection between the ranges of the first bounding box and the second bounding box, then determine that the text embedding position score of the embedded text is a preset value.

[0018] Optionally, for the method as above, obtaining the reasonableness of the embedded text being on the entity corresponding to the entity category by interacting with the large language model according to the entity category corresponding to the second bounding box containing the first bounding box includes:

[0019] Determine the depicted scene of the target image;

[0020] According to the entity category corresponding to the second bounding box containing the first bounding box, obtain the reasonableness of the embedded text being on the entity corresponding to the entity category under the depicted scene by interacting with the large language model.

[0021] Optionally, for the method as above, determining the first bounding box of the embedded text and the second bounding box of the non-text entities in the target image includes:

[0022] Input the target image into the vision large model for object recognition to obtain the first bounding box of the embedded text and the second bounding box of the recognized non-text entities;

[0023] Or, use the optical character recognition algorithm to perform text detection on the target image to obtain the first bounding box of the embedded text.

[0024] Optionally, for the method as above, determining the text layout score of the embedded text based on the target image includes:

[0025] Input the target image into the text layout evaluation model to obtain a text arrangement score, a text size score, and a text angle score; wherein, the text arrangement score represents the positional rationality and arrangement aesthetics of the text arrangement; the text size score represents the degree of consistency in the size of the embedded text; the text angle score represents the degree of rationality in the angle of the embedded text.

[0026] Obtain the average value of the text arrangement score, the text size score, and the text angle score as the text layout score of the embedded text.

[0027] Optionally, in the above method, the target image contains multiple embedded texts. Based on the target image, according to the pre-trained text aesthetics evaluation model, obtain the aesthetics score of the embedded text, including:

[0028] Input the target image into the pre-trained text aesthetics evaluation model to obtain the aesthetics score of each embedded text.

[0029] Obtain the average value of the aesthetics scores of multiple embedded texts as the aesthetics score of the embedded text.

[0030] In a second aspect, the present application provides an image embedded text quality evaluation device, including:

[0031] An acquisition module for acquiring a target image, where the target image is an image containing embedded text;

[0032] A first obtaining module for obtaining the aesthetics score of the embedded text based on the target image according to the pre-trained text aesthetics evaluation model;

[0033] A second obtaining module for obtaining the text embedding position score of the embedded text according to the positional relationship between the embedded text and the entity in the target image;

[0034] A determination module for determining the text layout score of the embedded text based on the target image;

[0035] A third obtaining module for performing weighted summation processing on the aesthetics score, the text embedding position score, and the text layout score of the embedded text to obtain the quality evaluation result of the embedded text.

[0036] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor;

[0037] The memory stores computer execution instructions;

[0038] The processor executes the computer execution instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementation manners of the first aspect.

[0039] Fourthly, an embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the above first aspect and / or various possible implementation manners of the first aspect when executed by a processor.

[0040] Fifthly, an embodiment of the present application provides a computer program product including a computer program, which implements the above first aspect and / or various possible implementation manners of the first aspect when executed by a processor.

[0041] A method, apparatus, and storage medium for evaluating the quality of image-embedded text provided by the present application obtain a target image, where the target image is an image containing embedded text; based on the target image, obtain the aesthetic score of the embedded text according to a pre-trained text aesthetic evaluation model; obtain the text embedding position score of the embedded text according to the positional relationship between the embedded text and the entities in the target image; determine the text layout score of the embedded text based on the target image; and perform weighted summation processing on the aesthetic score, text embedding position score, and text layout score of the embedded text to obtain the quality evaluation result of the embedded text. By comprehensively considering multiple aspects such as the aesthetic evaluation of the embedded text, text embedding position, and text layout score for quality evaluation, it not only evaluates the quality of the text itself (text aesthetic evaluation), but also includes the relationship between the text and the image (text embedding position evaluation, text layout), thereby making the quality evaluation result of the embedded text more accurate and comprehensive. Description of the Drawings

[0042] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.

[0043] Figure 1 It is a schematic flowchart of a method for evaluating the quality of image-embedded text provided by an embodiment of the present application;

[0044] Figure 2 It is a schematic flowchart of another method for evaluating the quality of image-embedded text provided by an embodiment of the present application;

[0045] Figure 3 It is a schematic structural diagram of an apparatus for evaluating the quality of image-embedded text provided by an embodiment of the present application;

[0046] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application.

[0047] Through the above-mentioned accompanying drawings, specific embodiments of the present application have been shown, and will be described in more detail hereinafter. These drawings and the written description are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by reference to specific embodiments. Detailed Description of Specific Embodiments

[0048] Here, exemplary embodiments will be described in detail, and examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0049] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards, and corresponding operation entrances are provided for the user to choose to authorize or refuse.

[0050] Currently, existing text evaluation methods only evaluate the text itself, such as the quality of a single font. However, directly applying existing text evaluation methods to evaluate embedded text in images can only evaluate the text itself, resulting in inaccurate evaluation results.

[0051] A method, device, and storage medium for evaluating the quality of embedded text in images provided by the present application comprehensively consider the aesthetics of the embedded text, the positional relationship between the embedded text and non-text entities, and the text layout of the embedded text, thereby obtaining an evaluation result of the quality of the embedded text. By evaluating not only the text itself but also the positional relationship between the embedded text and the image, the evaluation is more targeted at the characteristics of the embedded text, enabling a comprehensive and accurate evaluation of the image with embedded text.

[0052] The technical solution of the present application and how the technical solution of the present application solves the above technical problems will be described in detail below with specific embodiments. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.

[0053] Figure 1 It is a schematic flowchart of a method for evaluating the quality of embedded text in an image provided by an embodiment of the present application. As Figure 1 shown, the method includes:

[0054] S101. Obtain a target image, where the target image is an image containing embedded text.

[0055] Among them, the target image can be generated by a text-to-image model or designed by a person, but it needs to contain embedded text.

[0056] S102. Based on the target image, according to a pre-trained text aesthetics evaluation model, obtain the aesthetics score of the embedded text.

[0057] Among them, the pre-trained text aesthetics evaluation model is used to score the aesthetics of the embedded text in the target image.

[0058] Specifically, input the target image into the pre-trained text aesthetics evaluation model to obtain the aesthetics score of the embedded text. The aesthetics score can be the direct output of the model or the result of further processing of the model output.

[0059] S103. According to the positional relationship between the embedded text and the entities in the target image, obtain the text embedding position score of the embedded text.

[0060] Among them, by evaluating the positional relationship between the embedded text and the entities in the target image, the text embedding position score of the embedded text is obtained.

[0061] Specifically, the rationality of the positional relationship can be evaluated to obtain the corresponding text embedding position score of the embedded text.

[0062] In the embodiments of the present application, obtaining the text embedding position score of the embedded text according to the positional relationship between the embedded text and the entities in the target image includes:

[0063] According to the target image, determine the first bounding box of the embedded text and the second bounding box of the non-text entities in the target image;

[0064] According to the positional relationship between the first bounding box and the second bounding box, obtain the text embedding position score of the embedded text.

[0065] Among them, the first bounding box is the bounding box of the embedded text, and there can be one or more in the target image.

[0066] The second bounding box is the second bounding box of the non-text entities, specifically, it can be the bounding box of a person or an object.

[0067] The positional relationship includes: the second bounding box completely contains the first bounding box, the second bounding box and the first bounding box have no overlapping part, and the first bounding box and the second bounding box have an overlapping part.

[0068] Obtain the bounding box of the embedded text in the target image as the first bounding box, and the bounding box of non-text entities (such as people, objects, etc.) as the second bounding box. The data format of the first bounding box is as follows: [(text, BBOX 文字 ), and the data format of the second bounding box is as follows: (category of entity 1, BBOX 实体1 ), (category of entity 2, BBOX 实体2 ), … BBOX refers to the bounding box. When the bounding box is a rectangle, the BBOX data contains the positions of the four vertices of the bounding box, which are the positions of the corresponding bounding box. When the bounding box is an irregular image, the BBOX data contains the positions of multiple points on the bounding box.

[0069] Based on the positions of the bounding box of the embedded text and the bounding box of the entity, the positional relationship between the bounding box of the embedded text and the bounding box of the entity can be confirmed. The positional relationship is used to confirm whether the text is reasonably embedded in the image, so as to obtain the corresponding embedded text embedding position score according to the positional relationship.

[0070] The advantage of such a setting is that by calculating the positional relationship between the text and other entities in the image, it is possible to more accurately identify whether the text is reasonably embedded in the image.

[0071] In the embodiments of the present application, obtaining the text embedding position score of the embedded text according to the positional relationship between the first bounding box and the second bounding box includes:

[0072] If the first bounding box is within the range of the second bounding box, then according to the entity category corresponding to the second bounding box containing the first bounding box, the reasonableness of the embedded text on the entity corresponding to the entity category is obtained through interaction with the large language model, and the text embedding position score of the embedded text is determined according to the reasonableness;

[0073] If the first bounding box is outside the range of the second bounding box, then the reasonableness of the embedded text on the non-entity is obtained through interaction with the large language model, and the text embedding position score of the embedded text is determined according to the reasonableness;

[0074] If there is an intersection between the ranges of the first bounding box and the second bounding box, then determine that the text embedding position score of the embedded text is a preset value.

[0075] Among them, the entity category of the second bounding box can be any category, such as face, fan, painting, book, cup, pen holder, T-shirt, tree, flower, ground, etc.

[0076] The preset value can be preset by the user according to the situation and can be any value. However, generally, there are relatively large problems when there is an intersection between the text and the entities in the image. Therefore, the preset values are all relatively small. Specifically, it can be 0.

[0077] If the area within the text bounding box (i.e., the first bounding box) is A, and the area within the bounding box of entity i (i.e., the second bounding box) is Bi, where i = 1, 2, …, n, and n is the number of entities. If the text bounding box (i.e., the first bounding box) is located within any one of the entity bounding boxes (i.e., the second bounding boxes), that is, all points in area A belong to area Bi (1 ≤ i ≤ n), it is expressed as It means that the text is embedded on this entity, which is called the first positional relationship. For the first positional relationship, by asking the LLM (Large Language Model) interactively whether it is reasonable for the text to be on an entity of the corresponding entity category, the specific score of the first positional relationship can be obtained. For example, when the entity category is an entity that itself contains fonts, the degree of reasonableness is relatively high. Specifically: when the text is on the surface of a fan, a painting, a book, a cup, a pen holder, a T-shirt, etc., it is considered reasonable, so the text embedding score is relatively high; when the entity category is an entity that does not contain fonts itself, the degree of reasonableness is relatively low. Specifically: when the text is on the surface of a human face, a tree, a flower, the ground, etc., it is considered less reasonable, so the text embedding score is relatively low.

[0078] If the text bounding box (i.e., the first bounding box) is located outside all entity bounding boxes (i.e., the second bounding boxes), that is, area A and area have no common points, it is expressed as where represents the empty set, which means that the text is in the blank or white space area of the image, called the second positional relationship. Then, by asking the LLM interactively about the reasonableness of the embedded text being on a non-entity, the embedding position score of the embedded text can be obtained. Specifically, for example, if a poem or several lines of a poem are in the blank space of the image, the possibility of this situation occurring in reality is relatively high, so the degree of reasonableness is relatively high, and a score above 90 will be obtained. If a trademark name is in the blank space of the image, the possibility of this situation occurring in reality is relatively low, and it generally will be embedded in an entity, so a relatively low score below 30 will be obtained.

[0079] If the text bounding box (i.e., the first bounding box) intersects with a single entity bounding box (i.e., the second bounding box), that is, area A and Bi (1 ≤ i ≤ n) have both common points and different points, and A and Bi partially overlap, it is expressed as and It means that the text has overflowed the entity range, called the third positional relationship; if the bounding box of the text intersects with the bounding boxes of two or more entities, that is, area A and areas Bi and Bj all have common points, and i ≠ j, it is expressed as A and If \(i\neq j\), it means that the text spans two different entities, which is called the fourth positional relationship. For both the third and fourth positional relationships where the ranges of the first bounding box and the second bounding box have an intersection, the text embedding position score of the embedded text is 0.

[0080] The advantages of this setting are as follows: By querying the LLM, the system can obtain a rationality assessment of the text embedding on a specific entity category. This semantic-level understanding goes beyond simple geometric positional relationships, making the embedding position score more intelligent and user-friendly. By automatically calculating the embedding position score, the system can optimize the text embedding position without manual intervention. This automation ability is particularly important in large-scale image processing and real-time applications. By combining positional relationships and semantic understanding, an intelligent text embedding position evaluation mechanism is provided, significantly improving the effect and efficiency of image processing.

[0081] In the embodiments of the present application, according to the entity category corresponding to the second bounding box containing the first bounding box, by interacting with the large language model through querying, the rationality degree of the embedded text being on the entity corresponding to the entity category is obtained, including:

[0082] Determine the depicted scene of the target image;

[0083] According to the entity category corresponding to the second bounding box containing the first bounding box, by interacting with the large language model through querying, obtain the rationality degree of the embedded text being on the entity corresponding to the entity category in the depicted scene.

[0084] Among them, the depicted scene can be a life or work scene, or an activity scene, such as: a concert, a ball game, or other activity scenes.

[0085] The depicted scene will affect the rationality degree of the embedded text being on the entity corresponding to the entity category. Different scenes result in different rationality degrees of the embedded text being on the entity corresponding to the entity category. Specifically, in life and work scenes, it is unreasonable to have text on a person's face, and the rationality degree is relatively low. However, in a game scene, a concert, etc., it is reasonable to have a small amount of text on a person's face, and the rationality degree is relatively high. In a commercial scene, text may be more suitable to appear on a billboard, while in a natural scenery, text may be more suitable to appear on a sign.

[0086] Therefore, it is necessary to determine the depicted scene of the target image, and then obtain the rationality degree of the embedded text being on the entity corresponding to the entity category in the depicted scene by interacting with the LLM through querying.

[0087] The advantage of this setting is that by recognizing the depicted scene of the target image, the system can understand the rationality of text embedding in a broader context. The introduction of the LLM enables the system to perform more complex reasoning and judgment based on semantic information. By querying the LLM, the system can obtain rationality suggestions regarding specific scenes and entity categories, thereby improving the accuracy of evaluation.

[0088] In the embodiment of the present application, determining the first bounding box of the embedded text and the second bounding box of the non-text entity in the target image includes:

[0089] Inputting the target image into a vision large model for object recognition to obtain the first bounding box of the embedded text and the second bounding box of the recognized non-text entity;

[0090] Alternatively, using an optical character recognition algorithm to perform text detection on the target image to obtain the first bounding box of the embedded text.

[0091] Among them, the vision large model is used to recognize the embedded text, obtain the first bounding box of the embedded text, and recognize non-text entities to obtain the second bounding box of the non-text entity.

[0092] The deep learning OCR recognition method mainly consists of two steps currently:

[0093] (1) Text detection in the target picture, that is, framing the text in the picture with a text box.

[0094] (2) Recognizing the text in the text box as the first bounding box of the embedded text.

[0095] The advantage of this setting is that by combining the vision large model and OCR technology, it realizes the efficient recognition of embedded text and non-text entities in the image and the determination of the bounding box, providing a solid foundation for subsequent image processing and analysis.

[0096] S104. Based on the target image, determine the text layout score of the embedded text.

[0097] Among them, the layout score of the embedded text can be obtained by inputting the target image into a text layout evaluation model. The text layout evaluation model is a neural network model trained with supervised data, mainly for evaluating the arrangement layout of the embedded text in the image.

[0098] S105. Perform weighted summation processing on the aesthetics score, text embedding position score, and text layout score of the embedded text to obtain the quality evaluation result of the embedded text.

[0099] Among them, the quality evaluation result of the embedded text includes sub-item scores and a final score. The sub-item scores are evaluated and analyzed from three aspects: the aesthetics of the text, the embedding position of the text, and the text layout.

[0100] The final score P is:

[0101] P = α × I + β × Q + δ × C

[0102] Among them, I is the score for the aesthetics of the text, Q is the score for the embedding position of the text, C is the score for the text layout, and α, β, and δ are the weights of the sub-item scores. Each weight can be preset by the user according to requirements.

[0103] Sub-item scores are provided and evaluated and analyzed from three aspects: the aesthetics of the text, the embedding position of the text, and the text layout. The analysis of the aesthetics of the text is obtained according to the output of a pre-trained text aesthetics evaluation model. The output of the model includes the position and aesthetics score of each text in the image. Therefore, the texts with lower aesthetics scores and their locations can be analyzed, described in the report, and boxed and marked in the image. The analysis of the embedding position of the text is obtained according to the positional relationship between the embedded text and the entities in the target image, including the specific type of positional relationship, the entity category of the corresponding entity where the embedded text is located, and the reason for the rationality of the positional relationship obtained by combining the LLM analysis (such as "it is relatively rare to embed text on the ground, and it does not conform to human aesthetic requirements and visual habits" etc.). The analysis of the text layout is obtained according to the output of the text layout evaluation model, including whether the arrangement of the embedded text is reasonable, whether the setting of the text size is appropriate, and whether it is coordinated with the setting of the text angle.

[0104] A method for evaluating the quality of image-embedded text provided by the present application includes obtaining a target image, where the target image is an image containing embedded text; based on the target image, obtaining the aesthetics score of the embedded text according to a pre-trained text aesthetics evaluation model; obtaining the text embedding position score of the embedded text according to the positional relationship between the embedded text and the entities in the target image; determining the text layout score of the embedded text based on the target image; and performing weighted summation processing on the aesthetics score of the embedded text, the text embedding position score, and the text layout score to obtain the quality evaluation result of the embedded text. By comprehensively considering multiple aspects such as the aesthetics evaluation of the embedded text, the text embedding position, and the text layout score for quality evaluation, not only the quality of the text itself (text aesthetics evaluation) is evaluated, but also the relationship between the text and the image (text embedding position evaluation, text layout) is included, thereby making the quality evaluation result of the embedded text more accurate and comprehensive.

[0105] In the above embodiment of the present application, determining the text layout score of the embedded text based on the target image includes:

[0106] The target image is input into the text layout evaluation model to obtain the text arrangement score, text size score, and text angle score; the text arrangement score represents the rationality of the position of the text arrangement and the aesthetics of the arrangement; the text size score represents the consistency of the size of the embedded text; and the text angle score represents the rationality of the angle of the embedded text.

[0107] The average value of the text arrangement score, the text size score, and the text angle score is obtained as the text layout score of the embedded text.

[0108] The text layout evaluation model is a neural network model trained with supervised data, which mainly evaluates the arrangement and layout of embedded text in the target image. The text layout evaluation model outputs three scores according to the input target image.

[0109] The training data for training the text layout evaluation model includes training images, which contain text bounding boxes and scoring data. The embedded text scores of the training images are completed by experts, and one image corresponds to three scores (the scores are in percentage), namely text arrangement score, text size score and text angle score.

[0110] There are two sources of images. One is images with higher standards obtained from the Internet or private databases, and the other is images obtained by processing better images. The processing methods include changing the position of text, the relative size of text, the angle of text, disrupting the order of text, etc.

[0111] The text sorting score refers to whether the arrangement of the text is reasonable, such as whether the adjacent text is in close position, whether the text is beautifully arranged, etc. The text size score refers to whether the size of the text in the image is reasonable, whether the key points are highlighted, whether the overall situation is coordinated, etc. It mainly targets the situation where the text size in the image is inconsistent. The text angle score refers to whether the angle of the text in the image is reasonable, such as whether the tilt angle is appropriate, whether the overall situation is coordinated, etc. When scoring the text position, text size and text tilt angle, the semantic meaning should be considered. For example, if the four characters "万事如意" are arranged in a semantic order, the position score will be high. If they are arranged in a disordered order, such as "万如意事", then its text arrangement score will definitely be low. Similarly, the text size and text angle need to consider semantics.

[0112] Get the average of the text arrangement score, text size score, and text angle score to get the text layout score:

[0113]

[0114] Among them, the text arrangement score is C1, the text size score is C2, and the text angle score is C3.

[0115] The advantages of such a setting are as follows: By automatically calculating the text arrangement, size, and angle scores through the model, the subjectivity and inconsistency of manual evaluation are reduced, and the efficiency and accuracy of evaluation are improved. By evaluating the text arrangement, size, and angle separately, the layout quality of the embedded text can be comprehensively analyzed, ensuring the beauty and readability of the text from multiple dimensions. The text arrangement score can help identify and optimize the position and arrangement of the text in the image to ensure its rationality and beauty.

[0116] In the above embodiments of the present application, the target image contains multiple embedded texts. Based on the target image, according to the pre-trained text beauty evaluation model, the beauty score of the embedded text is obtained, including:

[0117] Input the target image into the pre-trained text beauty evaluation model to obtain the beauty score of each embedded text;

[0118] Obtain the average value of the beauty scores of multiple embedded texts as the beauty score of the embedded text.

[0119] Among them, the training data of the pre-trained text beauty evaluation model includes a large number of images of embedded texts, the positions of the texts in the images (text bounding boxes, represented as a 4D array, which are the coordinates of the upper left corner, upper right corner, lower right corner, and lower left corner of the bounding box), and the beauty scores (percentage scores) of each text. After iterative optimization of the initial neural network, a pre-trained text beauty evaluation model is obtained. The acquisition of training data includes two ways. One is to collect images with higher beauty scores containing embedded texts from public networks and private databases, and the other is the images obtained after processing these images with higher beauty scores. The processing methods include changing the colors and brightness of some texts, and distorting and deforming some texts, etc. The beauty score of the directly collected image text is 100 full marks, and the beauty score of the text in the processed image is manually scored by experts.

[0120] Input the image containing embedded texts into the pre-trained text beauty evaluation model. The model outputs the position of each text and the beauty score of each text. Finally, the beauty score of the embedded text is obtained by taking the average value of the beauty scores of all texts:

[0121]

[0122] Among them, I is the overall text beauty score of the image, n is the number of texts in the image, and I i is the beauty score of the i-th text output by the text beauty evaluation model.

[0123] The advantage of this setting is that through this systematic aesthetic evaluation method, designers can better control and optimize the visual effects of the embedded text, improving the overall aesthetics and user satisfaction of the final work.

[0124] Figure 2 Another quality evaluation method for image-embedded text provided by the embodiments of this application is executed through a quality evaluation system for image-embedded text. The quality evaluation system for image-embedded text includes a test text generation module, a test image generation module, a text quality evaluation module, and a test result analysis module, including:

[0125] S201. The test text generation module automatically generates test texts.

[0126] Among them, several themes are built into the test text generation module, such as poems, product promotions, dates, image themes, signatures, inscriptions, etc. The LLM (Large Language Model) will generate several test texts for each theme in turn. The test texts include 3 levels of difficulty: simple, medium, and difficult. Therefore, for each text-to-image generation large model to be tested, a total of 3*m*n*q test texts are generated, where m is the number of test samples included in each theme, n is the number of themes, and q is the number of font styles. The generation process is automatically completed, and the test texts for different text-to-image generation large models are different.

[0127] The LLM generates test texts according to the theme, font, and difficulty level. For example, the LLM inputs the command "Generate a test text for the text-to-image generation large model, which includes a quatrain, and specify the font as regular script, and the difficulty of text-to-image is difficult". The test text generated by the LLM is "Please use the beauty of regular script calligraphy to create an image that integrates poetry. In the image, every stroke reveals the charm of regular script, slowly outlining the quatrain of the Tang Dynasty poet Wang Zhihuan: 'The sun along the mountain bows; The Yellow River seawards flows. You can enjoy a great sight; By climbing to a greater height.' Let this work not only show the elegance of regular script, but also make the viewer seem to be able to look far and wide with the poem, and the heart is suddenly enlightened."

[0128] S202. The test image generation module is responsible for inputting the test text into the text-to-image generation large model to be evaluated and outputting the test image.

[0129] Among them, the test text is input into the text-to-image generation large model to be evaluated in turn, and the image output by the text-to-image generation large model is used as the test image, which forms (text, image) test image data with the test text. These test image data are the data for subsequent evaluation. This module includes multiple evaluation items, namely text alignment evaluation, text aesthetics evaluation, text embedding position evaluation, and text layout evaluation.

[0130] S203. The text quality assessment module is responsible for assessing the quality of the text in the test image data. Among them, the text quality assessment module includes multiple evaluation items, namely text alignment evaluation, text aesthetics evaluation, text embedding position evaluation, and text layout evaluation.

[0131] Among them, for text alignment evaluation, it mainly examines whether the test image contains the text required to be embedded in the test text; for text aesthetics evaluation, it mainly examines the aesthetics of individual characters generated in the test image; for text embedding position, it mainly examines whether the position of the text in the test image is appropriate (if the test text clearly specifies the position where the text should be embedded, this item checks whether the text position is consistent with the required position. If the test text does not clearly specify the position of the text, this item determines whether the position of the text in the test image is appropriate based on the overall aesthetics of the image and the habitual position of text embedding); for text layout, it mainly examines whether the arrangement of the overall text and the size and inclination angle of each character are reasonable.

[0132] The text alignment evaluation is as follows: Input the test text into the large language model LLM to identify the text content required to be embedded in the test image, and then input the test image into the character recognition algorithm to detect all the text contained in the output image. Sequentially query whether the text required to be embedded in the test text is already included in the text in the image, and calculate the percentage of text coverage as the text matching degree.

[0133] The text aesthetics evaluation is as follows: Input the test image into the text aesthetics evaluation model to output the text aesthetics score.

[0134] The text embedding position evaluation is as follows: If the position where the text needs to be embedded in the image is specified, it will be judged whether the position of the text bounding box is consistent with the required position. If the position where the text needs to be embedded in the image is not specified, input the test image into the vision large model to obtain the text bounding box and the entity bounding box in the image, and then obtain the text embedding position score through the position relationship between the text bounding box and the entity bounding box in the image.

[0135] The text layout evaluation is as follows: Input the test image into the text evaluation model to obtain the text arrangement score, text size score, and text angle score, and obtain the text layout score based on the text arrangement score, text size score, and text angle score.

[0136] S204. The evaluation result analysis module synthesizes multiple evaluation items of the text quality assessment module, analyzes the evaluation items, and gives an evaluation report.

[0137] The evaluation result analysis module analyzes by integrating the results of each evaluation item in the LLM comprehensive text quality evaluation. Fill in the report content according to the evaluation report template, including the overall evaluation and individual evaluations of text alignment, text aesthetics, text embedding position, and text layout, and give improvement suggestions.

[0138] Another method for evaluating the quality of text embedded in images provided by the embodiments of the present application can fully automate the evaluation process without manual intervention; it does not require (text, image) pairs as test samples, that is, it does not require the ground truth images corresponding to the test texts; it does not require standard fonts for comparison, and directly uses a neural network model to evaluate the aesthetics of individual characters.

[0139] Figure 3 FIG. is a structural example diagram of a device for evaluating the quality of text embedded in images provided by the embodiments of the present application. As shown in the figure, the device 30 for evaluating the quality of text embedded in images includes: an acquisition module 301, a first obtaining module 302, a second obtaining module 303, a determination module 304, and a third obtaining module 305. Among them:

[0140] The acquisition module 301 is used to acquire a target image, where the target image is an image containing embedded text;

[0141] The first obtaining module 302 is used to obtain the aesthetics score of the embedded text based on the target image according to a pre-trained text aesthetics evaluation model;

[0142] The second obtaining module 303 is used to obtain the text embedding position score of the embedded text according to the positional relationship between the embedded text and the entities in the target image;

[0143] The determination module 304 is used to determine the text layout score of the embedded text based on the target image;

[0144] The third obtaining module 305 is used to perform a weighted summation process on the aesthetics score, text embedding position score, and text layout score of the embedded text to obtain the quality evaluation result of the embedded text.

[0145] In a possible implementation manner, the acquisition module 301 is used for:

[0146] Input the target image into a pre-trained text aesthetics evaluation model to obtain the aesthetics score of each embedded text;

[0147] Obtain the average value of the aesthetics scores of multiple embedded texts as the aesthetics score of the embedded text.

[0148] In a possible implementation manner, the second obtaining module 303 is used for:

[0149] Determine a first bounding box of the embedded text and a second bounding box of non-text entities in the target image according to the target image;

[0150] Obtain a text embedding position score of the embedded text according to the positional relationship between the first bounding box and the second bounding box.

[0151] In a possible implementation, the second obtaining module 303 is further configured to:

[0152] If the first bounding box is within the range of the second bounding box, then according to the entity category corresponding to the second bounding box containing the first bounding box, obtain, through interaction with a large language model, the reasonableness of the embedded text being on an entity of the corresponding entity category, and determine the text embedding position score of the embedded text according to the reasonableness;

[0153] If the first bounding box is outside the range of the second bounding box, then obtain, through interaction with a large language model, the reasonableness of the embedded text being on a non-entity, and determine the text embedding position score of the embedded text according to the reasonableness;

[0154] If there is an intersection between the ranges of the first bounding box and the second bounding box, then determine that the text embedding position score of the embedded text is a preset value.

[0155] In a possible implementation, the second obtaining module 303 is further configured to:

[0156] Determine the depicted scene of the target image;

[0157] According to the entity category corresponding to the second bounding box containing the first bounding box, obtain, through interaction with a large language model, the reasonableness of the embedded text being on an entity of the corresponding entity category under the depicted scene.

[0158] In a possible implementation, the second obtaining module 303 is further configured to:

[0159] Input the target image into a vision large model for object recognition to obtain a first bounding box of the embedded text and a second bounding box of the recognized non-text entities;

[0160] Alternatively, use an optical character recognition algorithm to perform text detection on the target image to obtain a first bounding box of the embedded text.

[0161] In a possible implementation, the determining module 304 is configured to:

[0162] Input the target image into a text layout evaluation model to obtain a text arrangement score, a text size score, and a text angle score; wherein, the text arrangement score represents the positional reasonableness and arrangement aesthetics of the text arrangement; the text size score represents the degree of consistency of the sizes of the embedded text; the text angle score represents the reasonableness of the angle of the embedded text;

[0163] Obtain the average value of the text arrangement score, the text size score, and the text angle score as the text layout score of the embedded text.

[0164] The device provided by the embodiments of the present application can execute the technical solutions shown in the above method embodiments, and the implementation principles and beneficial effects are similar, so they will not be elaborated here.

[0165] It should be noted that it should be understood that the division of each module of the above device is only a logical function division. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by a processing element; they can also all be implemented in hardware; they can also be partially implemented in the form of software called by a processing element and partially implemented in hardware. For example, the processing module can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and called and executed by a certain processing element of the above device to perform the functions of the above processing module. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together or can be independently implemented. Here, the processing element can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the processor element or the instruction in software form.

[0166] For example, the above modules can be one or more integrated circuits configured to implement the above method, such as: one or more Application Specific Integrated Circuits (ASICs), or, one or more Digital Signal Processors (DSPs), or, one or more Field Programmable Gate Arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a processing element scheduling program code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processors that can call program code. Again, these modules can be integrated together and implemented in the form of a System-On-a-Chip (SOC).

[0167] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a Digital Video Disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0168] Figure 4 Schematic structural diagram of the electronic device provided in the embodiments of the present application. As Figure 4 shown, the electronic device 40 includes:

[0169] The electronic device 40 may include a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a communication component 403, and other components. Among them, the processor 401, the memory 402, and the communication component 403 are connected through a bus 404.

[0170] In a specific implementation process, at least one processor 401 executes the computer execution instructions stored in the memory 402, so that at least one processor 401 executes the above method for evaluating the quality of embedded text in an image.

[0171] For the specific implementation process of the processor 401, reference can be made to the above method embodiments, and the implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.

[0172] In the above Figure 4In the illustrated embodiments, it should be understood that the processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in connection with the invention may be directly embodied as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.

[0173] The memory may include random access memory (RAM), and may also include non-volatile memory (NVM), such as at least one disk memory.

[0174] The bus may be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the buses in the drawings of the present application are not limited to only one bus or one type of bus.

[0175] In some embodiments, a computer program product is also proposed, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps in any of the above-mentioned methods for evaluating the quality of inlaid text in an image are implemented.

[0176] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, and details will not be repeated here.

[0177] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling relevant hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0178] To this end, the embodiments of the present application provide a computer-readable storage medium, which stores multiple instructions that can be loaded by a processor to execute the steps in any of the methods for evaluating the quality of inlaid text in an image provided by the embodiments of the present application.

[0179] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.

[0180] According to one aspect of the present application, there is provided a computer program product or a computer program, and the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium.

[0181] Since the instructions stored in the storage medium can execute the steps in any of the image-embedded text quality evaluation methods provided in the embodiments of the present application, the beneficial effects achievable by any of the image-embedded text quality evaluation methods provided in the embodiments of the present application can be achieved. For details, see the foregoing embodiments and will not be elaborated here.

[0182] Those skilled in the art will readily think of other implementations of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and the embodiments are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.

[0183] It should be understood that the present application is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A method for evaluating the quality of text embedded in an image, characterized in that Including: Obtain a target image, where the target image is an image containing embedded text; Based on the target image, obtain the aesthetics score of the embedded text according to a pre-trained text aesthetics evaluation model; According to the positional relationship between the embedded text and the entities in the target image, obtain the text embedding position score of the embedded text; Based on the target image, determine the text layout score of the embedded text; Perform weighted summation processing on the aesthetics score of the embedded text, the text embedding position score, and the text layout score to obtain the quality evaluation result of the embedded text.

2. The method according to claim 1, characterized in that The step of obtaining the text embedding position score of the embedded text according to the positional relationship between the embedded text and the entities in the target image includes: According to the target image, determine the first bounding box of the embedded text and the second bounding box of the non-text entities in the target image; According to the positional relationship between the first bounding box and the second bounding box, obtain the text embedding position score of the embedded text.

3. The method according to claim 2, wherein The step of obtaining the text embedding position score of the embedded text according to the positional relationship between the first bounding box and the second bounding box includes: If the first bounding box is within the range of the second bounding box, according to the entity category corresponding to the second bounding box containing the first bounding box, obtain the reasonableness of the embedded text being on the entity corresponding to the entity category by interacting with a large language model, and determine the text embedding position score of the embedded text according to the reasonableness; If the first bounding box is outside the range of the second bounding box, obtain the reasonableness of the embedded text being on a non-entity by interacting with a large language model, and determine the text embedding position score of the embedded text according to the reasonableness; If there is an intersection between the ranges of the first bounding box and the second bounding box, determine that the text embedding position score of the embedded text is a preset value.

4. The method according to claim 3, wherein The step of obtaining the reasonableness of the embedded text being on the entity corresponding to the entity category by interacting with a large language model according to the entity category corresponding to the second bounding box containing the first bounding box includes: Determine the depicted scene of the target image; According to the entity category corresponding to the second bounding box containing the first bounding box, obtain the reasonableness of the embedded text being on the entity corresponding to the entity category in the depicted scene by interacting with a large language model.

5. The method according to claim 2, characterized in that, The step of determining the first bounding box of the embedded text and the second bounding box of the non-text entities in the target image according to the target image includes: Input the target image into a vision large model for object recognition to obtain the first bounding box of the embedded text and the second bounding box of the recognized non-text entities; Alternatively, use an optical character recognition algorithm to perform text detection on the target image to obtain the first bounding box of the embedded text.

6. The method according to any one of claims 1-5, characterized in that, The step of determining the text layout score of the embedded text based on the target image includes: Input the target image into a text layout evaluation model to obtain a text arrangement score, a text size score, and a text angle score; wherein, the text arrangement score represents the positional rationality and arrangement aesthetics of the text arrangement; the text size score represents the degree of consistency in the size of the embedded text; the text angle score represents the degree of rationality in the angle of the embedded text. Obtain the average value of the text arrangement score, the text size score, and the text angle score as the text layout score of the embedded text.

7. The method according to any one of claims 1-5, characterized in that, The target image contains multiple embedded texts. Based on the target image, according to a pre-trained text aesthetics evaluation model, obtaining the aesthetics score of the embedded text includes: Input the target image into a pre-trained text aesthetics evaluation model to obtain the aesthetics score of each embedded text. Obtain the average value of the aesthetics scores of the multiple embedded texts as the aesthetics score of the embedded text.

8. An image embedded text quality evaluation device, characterized in that Includes: An acquisition module for acquiring a target image, where the target image is an image containing embedded text. A first obtaining module for obtaining the aesthetics score of the embedded text based on the target image according to a pre-trained text aesthetics evaluation model. A second obtaining module for obtaining a text embedding position score of the embedded text according to the positional relationship between the embedded text and an entity in the target image. A determination module for determining the text layout score of the embedded text based on the target image. A third obtaining module for performing a weighted summation process on the aesthetics score of the embedded text, the text embedding position score, and the text layout score to obtain a quality evaluation result of the embedded text.

9. An electronic device, characterized in that, Includes: A memory, a processor; The memory stores computer execution instructions. The processor executes the computer execution instructions stored in the memory, so that the processor executes the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, Computer execution instructions are stored in the computer-readable storage medium, and when the computer execution instructions are executed, they are used to implement the method according to any one of claims 1-7.

11. A computer program product, characterized in that, Includes a computer program, and when the computer program is executed, it implements the method according to any one of claims 1-7.