Method, apparatus for generating image including multiple targets
Patent Information
- Application Number
- CN202510152789.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2026-08-18
AI Technical Summary
但是在将原始语言文字转换为包括更多目标的另一语言的文字时,可能由于该更多目标的布局风格与原始语言文字的布局风格不一致或该更多目标在原始语言文字的原始边界范围内排列拥挤或互相重叠而导致不能达到令人满意的显示效果
Smart Images

Figure CN122597555A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this disclosure relate to the field of image processing, and more specifically to methods and apparatus for generating images comprising multiple targets. Background Technology
[0002] Images often contain embedded text to reflect dialogue or other descriptive information (such as titles, onomatopoeia, scene descriptions, etc.) within a speech bubble. In some scenarios, it may be necessary to convert the original text in an image into another language to replace the original text, generating an image that includes the text in the other language. The text in the other language may contain more characters than the original text, i.e., more objects. However, when converting the original text to another language that includes more objects, unsatisfactory display results may not be achieved because the layout style of the additional objects is inconsistent with that of the original text, or because the additional objects are crowded or overlapped within the original boundaries of the original text. Summary of the Invention
[0003] The embodiments of this disclosure provide a method and apparatus for generating an image including multiple targets, which enables the layout style of the multiple targets in the generated image to be as consistent as possible with the original layout style, and eliminates the need for manual adjustment. Furthermore, some embodiments can further avoid the multiple targets being crowded or overlapping, thereby achieving efficient and optimized layout of multiple targets.
[0004] According to one aspect of this disclosure, at least one embodiment provides a method for generating an image including a plurality of targets, comprising: determining a second layout of a plurality of second targets to be included in a second image based on features of a first layout of a plurality of first targets included in a first image, wherein the number of the plurality of second targets is greater than the number of the plurality of first targets, and the difference or ratio between the range of the total bounding box including the plurality of second targets and the range of the total bounding box including the plurality of first targets is not greater than a predetermined threshold; generating a second image including the plurality of second targets based on the second layout and the first image, wherein determining the second layout of the plurality of second targets to be included in the second image based on features of the first layout of the plurality of first targets included in the first image comprises: determining at least a positional trend and / or a rotational angle trend of the plurality of first targets; and determining the position and / or rotational angle of each of the plurality of second targets based on the positional trend and / or rotational angle trend of the plurality of first targets.
[0005] According to another aspect of this disclosure, at least one embodiment provides an apparatus for generating an image comprising a plurality of targets, including components for performing various steps of a method according to at least one embodiment of this disclosure.
[0006] According to another aspect of this disclosure, at least one embodiment provides an apparatus for generating an image including a plurality of targets, comprising: a memory for storing computer instructions; and a processor for reading the computer instructions from the memory and performing a method according to at least one embodiment of this disclosure.
[0007] According to another aspect of this disclosure, at least one embodiment provides a non-transitory computer-readable storage medium having computer instructions stored thereon, wherein, when executed by a processor, the computer instructions cause the processor to perform a method according to at least one embodiment of this disclosure.
[0008] According to another aspect of this disclosure, at least one embodiment provides a computer program product including computer instructions, wherein, when executed by a processor, the computer instructions cause the processor to perform a method according to at least one embodiment of this disclosure. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments or related technologies of this disclosure, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 This demonstrates an example of translation failure in an application scenario where a Japanese text or a segment of text in an image is translated into another language to generate image content in that language for the user to read.
[0011] Figure 2 A schematic diagram of the hardware operating environment of an apparatus for running a method for generating an image comprising multiple targets according to at least one embodiment of the present disclosure is shown.
[0012] Figure 3 A flowchart is shown of a method for generating an image comprising multiple targets according to at least one embodiment of the present disclosure.
[0013] Figure 4A A schematic diagram illustrating an example of the steps for determining a second layout of a plurality of second targets to be included in a second image based on features of a first layout of a plurality of first targets included in a first image, according to at least one embodiment of the present disclosure.
[0014] Figure 4B A schematic diagram illustrates yet another example of the steps for determining a second layout of a plurality of second targets to be included in a second image based on features of a first layout of a plurality of first targets included in a first image, according to at least one embodiment of the present disclosure.
[0015] Figure 5A A schematic diagram illustrating another example of the steps for determining a second layout of a plurality of second targets to be included in a second image based on features of a first layout of a plurality of first targets included in a first image, according to at least one embodiment of the present disclosure.
[0016] Figure 5B A schematic diagram illustrating yet another example of the steps for determining a second layout of a plurality of second targets to be included in a second image based on features of a first layout of a plurality of first targets included in a first image, according to at least one embodiment of the present disclosure.
[0017] Figure 6 A schematic diagram is shown illustrating, according to at least one embodiment of the present disclosure, that the total bounding box of a plurality of second targets is limited to or is not significantly different from the total bounding box of a plurality of first targets.
[0018] Figure 7 This diagram illustrates a situation where excessive overlap between second targets leads to the generated second targets failing to display correctly or displaying incorrectly.
[0019] Figure 8 A schematic diagram is shown according to at least one embodiment of the present disclosure, in which the union of the bounding boxes of a plurality of second targets is within the total bounding box of the plurality of second targets and has the largest area, and the area of overlap between the bounding boxes of the plurality of second targets is minimized.
[0020] Figure 9 A schematic diagram is shown illustrating the use of a genetic algorithm according to at least one embodiment of the present disclosure to maximize the area of the union of the bounding boxes of a plurality of second targets within the total bounding box of the plurality of second targets, and minimize the area of overlap between the bounding boxes of the plurality of second targets. Detailed Implementation
[0021] Referring now to specific embodiments of this disclosure, examples of which are illustrated in the accompanying drawings. Although this application will be described in conjunction with specific embodiments, it will be understood that it is not intended to limit this application to the described embodiments. Rather, it is intended to cover variations, modifications, and equivalents included within the spirit and scope of this disclosure. It should be noted that the method steps described herein can be implemented by any functional block or functional arrangement, and any functional block or functional arrangement can be implemented as a physical entity or a logical entity, or a combination of both.
[0022] In this article, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0023] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0024] As previously explained, when converting text from an original language to another language that includes more targets, unsatisfactory display results may not be achieved due to inconsistencies in the layout style of the additional targets compared to the original language, or because the additional targets are crowded or overlapped within the original language's boundaries. For example, in applications that translate one or more Japanese characters from an image into English to generate English text for users to read, it is generally desirable that the translated dialogue or other descriptive information (such as titles, onomatopoeia, scene descriptions, etc.) within the original dialogue or other descriptive information's bounding box or slightly larger than the original bounding box. For instance, the translated dialogue or text should ideally be located within the original dialogue or text's bubble to prevent the translated text from overflowing the original bounding box and obscuring the original background image or other content that should not be obscured. However, the translated text often differs in character count from the original, especially when the translated text has more characters than the original. If a generative AI model is directly used to generate multiple translated targets (multiple characters of text) within the original bounding box, some targets may not display correctly, display incorrectly, be severely occluded, or have inconsistent layouts, resulting in an unpleasant visual effect. Such generation may fail or produce errors due to quality issues. On the other hand, manually arranging or adjusting multiple targets randomly or based on preference within the original bounding box leads to low efficiency; for example, manually adjusting 1000 images takes approximately 2.7 hours. Therefore, an efficient and optimized layout of multiple targets is needed.
[0025] Figure 1 This demonstrates an example of translation failure in an application scenario where a Japanese text or a segment of text in an image is translated into another language to generate image content in that language for the user to read.
[0026] like Figure 1As shown, the original image contains the Japanese character "ムシヤ" (an onomatopoeic word), which has three characters. This is just an example; other Japanese characters can also be used. Image processing techniques can extract the Japanese character from the image, and translation techniques can be used to translate it into the English word "MUSH." The translated "ムシヤ" into "MUSH" contains four characters (specifically, letters), more than the original three. If generative artificial intelligence is used to automatically generate the four target English characters "MUSH," some letters may be displayed incorrectly or severely occluded (e.g., the "U" may not be displayed) because the translated English text is expected to be roughly confined within the bounding box of the original Japanese text. Therefore, the quality of the generated "MUSH" is poor, leading to generation failure. Furthermore, the display layout of the English "MUSH" is inconsistent with the original Japanese "ムシヤ," resulting in an unharmonious visual effect. However, if people were to manually modify or determine the layout of the four target English words "MUSH", it would consume a lot of manpower and time.
[0027] The defects and problems existing in the above-mentioned prior art solutions are also the result of the inventor's careful research after practical and creative labor. The discovery process of the above problems and the solutions proposed by at least one embodiment disclosed below for the above problems are all creative contributions of the inventor during the invention process.
[0028] In view of this, embodiments of the present disclosure provide a technique for generating images comprising multiple targets, which enables the layout style of the multiple targets in the generated image to be as consistent as possible with the original layout style, and eliminates the need for manual adjustments. Furthermore, some embodiments can further avoid the multiple targets being crowded or overlapping, thereby achieving efficient and optimized layout of multiple targets.
[0029] Figure 2 A schematic diagram of the hardware operating environment for running an apparatus for generating an image comprising multiple targets according to at least one embodiment of the present disclosure is shown.
[0030] like Figure 2 As shown, the apparatus may include a processor 210 and a memory 220, the memory 220 being coupled to the processor 210 and storing computer instructions therein (e.g., a program for implementing a method for generating an image including a plurality of targets) for performing steps of various methods of at least one embodiment of the present disclosure when executed by the processor 210.
[0031] Processor 210 may include, but is not limited to, one or more processors or microprocessors.
[0032] The memory 220 may include, but is not limited to, random access memory (RAM), read-only memory (ROM), flash memory, EPROM memory, EEPROM memory, registers, computer storage media (e.g., hard disk, floppy disk, solid-state drive, removable disk, CD-ROM, DVD-ROM, Blu-ray disc, etc.).
[0033] In addition, the electronic device may also include (but is not limited to) a data bus 230, an input / output (I / O) bus 240, a display 250, and input / output devices 260 (e.g., a keyboard, mouse, speaker, etc.).
[0034] The processor 210 can communicate with external displays 250 and input / output devices 260 via the I / O bus 240.
[0035] Those skilled in the art will understand that Figure 2 The structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0036] Figure 3 A flowchart is shown of a method for generating an image comprising multiple targets according to at least one embodiment of the present disclosure.
[0037] like Figure 3 As shown, the method for generating an image that includes multiple targets includes steps 310 and 320.
[0038] In step 310, a second layout of multiple second targets to be included in the second image is determined based on the features of a first layout of multiple first targets included in the first image.
[0039] Here, the target can include various objects, such as people, animals, text, characters, etc. In this document, characters are used as examples to describe various embodiments, but this disclosure is not limited thereto. For example, using characters... Figure 1 For example, each first target can be a character of Japanese text, and each second target can be a character of English text translated from Japanese text.
[0040] Typically, the number of secondary targets is greater than the number of primary targets. Therefore, the secondary layout of the secondary targets to be included in the secondary image needs to be determined based on the characteristics of the primary layout of the primary targets included in the primary image. If the number of secondary targets is equal to or less than the number of primary targets, it is usually sufficient to place the secondary targets according to the primary layout of the primary targets, as this arrangement generally does not cause overlap or occlusion between targets and will not lead to failure.
[0041] Moreover, the difference or ratio between the range of the total bounding box including multiple second targets and the range of the total bounding box including multiple first targets is not greater than a predetermined threshold.
[0042] Here, the bounding box of a target means finding a region of a predetermined shape (e.g., a rectangular region or a circular region) that can completely contain all the pixels of the target (e.g., the region is the outer rectangle or outer circle of the target or slightly larger than the outer rectangle or outer circle of the target).
[0043] The extent of the total bounding box of multiple first targets (or second targets) (or referred to as the "joint bounding box", "outer rectangle", etc.) refers to finding a region of a predetermined shape (e.g., a rectangular region or a circular region) that can completely contain the bounding boxes of all individual targets (i.e., multiple first targets (or second targets)). (e.g., the region is the outer rectangle or outer circle of the multiple targets or is slightly larger than the outer rectangle or outer circle of the multiple targets.)
[0044] For example, a two-dimensional bounding box is typically represented by a quadruple (xmin, ymin, xmax, ymax), where (xmin, ymin) are the coordinates of the top-left corner of the bounding box, and (xmax, ymax) are the coordinates of the bottom-right corner. Taking a rectangle as an example, suppose the bounding boxes of two targets are B1=(x1min, y1min, x1max, y1max) and B2=(x2min, y2min, x2max, y2max). A new total bounding box (Bunion) can be found such that both B1 and B2 are completely contained within it, as follows.
[0045] Calculate the coordinates of the top-left corner of the new total bounding box:
[0046] xunion_min=min(x1min, x2min)
[0047] yunion_min=min(y1min, y2min)
[0048] Calculate the coordinates of the bottom right corner of the new total bounding box:
[0049] xunion_max = max(x1max, x2max)
[0050] yunion_max = max(y1max, y2max)
[0051] Combined into a new total bounding box (Bunion):
[0052] Bunion=(xunion_min, yunion_min, xunion_max, yunion_max).
[0053] Typically, it is desirable for the total bounding box containing multiple second targets to be less than or equal to the total bounding box containing multiple first targets, but it is also acceptable for it to be slightly larger. Here, for example, "slightly larger" could mean that the total bounding box containing multiple second targets is 1.05 times the total bounding box containing multiple first targets (i.e., the predetermined threshold could be 1.05 times, but of course, this value is not limited to this, as long as the newly generated multiple second targets do not cover the surrounding areas that should not have been covered in the first place).
[0054] For example, in translation applications, such as Figure 1 As shown, it is generally desirable that the translated dialogue or other descriptive information (such as titles, onomatopoeia, scene descriptions, etc.) in the speech bubble should also be confined within or slightly larger than the bounding box of the original dialogue or other descriptive information. For example, it is best that the translated text of one or more dialogues is located within the speech bubble of the original text of one or more dialogues, so as to prevent the translated text from overflowing the original bounding box and obscuring the original background image or other content that should not be obscured.
[0055] Here, the features of the first layout of the multiple first targets included in the first image may include at least positional trends and / or rotational angle trends, the position and extent (including length and width) of the total bounding box of the multiple first targets, etc. These objective features can reflect the layout rules and constraints of the multiple second targets to be included in the second image, thereby determining a suitable layout for the multiple second targets.
[0056] Step 310 may include: determining at least the position trend and / or rotation angle trend of a plurality of first targets; and determining the position and / or rotation angle of each of a plurality of second targets based on the position trend and / or rotation angle trend of the plurality of first targets.
[0057] As described above, the number of multiple second targets is usually greater than the number of multiple first targets, so it is necessary to determine how to arrange the second targets that outnumber the first targets.
[0058] Figure 4AA schematic diagram illustrating an example of step 310 of determining a second layout of a plurality of second targets to be included in a second image based on features of a first layout of a plurality of first targets included in a first image, according to at least one embodiment of the present disclosure. Step 310 of determining a second layout of a plurality of second targets to be included in a second image based on features of a first layout of a plurality of first targets included in a first image may include: determining positional trends of the plurality of first targets; and determining the position of each of the plurality of second targets based on the positional trends of the plurality of first targets.
[0059] In some embodiments, the multiple first targets may be multiple characters of a first language (e.g., Japanese), and the multiple second targets may be multiple characters of a second language (e.g., English), wherein the multiple characters of the second language are translations of the multiple characters of the first language, and the arrangement of the multiple first targets in the first layout and the arrangement of the multiple second targets in the second layout respectively conform to the language habits of the first language and the language habits of the second language. The language habits mentioned here refer to, for example, the word "MUSH" should be laid out as "MUSH" from left to right, or as follows if there are line breaks:
[0060] “MU”
[0061] "SH" is also acceptable.
[0062] The same applies to Japanese and other languages. This conforms to language conventions and at least allows people to understand its meaning in the language.
[0063] like Figure 4A As shown at the top, multiple primary targets are Japanese characters. "" to determine the Japanese characters " The positional trend of multiple first targets can be determined by: identifying the positional fit line of the center point of the bounding box of each first target as the positional trend of the multiple first targets. Figure 4A As shown, the Japanese characters " The center point of the bounding box of each character is identified, and these center points are fitted, which can include linear fitting to a straight line and non-linear fitting to a curve. Figure 4A In the example shown, these center points are linearly fitted to straight lines (represented by dashed lines with arrows), which serve as the positional trend of multiple primary targets.
[0064] So, from the Japanese characters " The translated English word "MUSH" should also follow the Japanese syllabary. The original position trend of "" is used to maintain the consistency of the text position before and after translation, making it easier for users to read.
[0065] like Figure 4A As shown in the lower part, based on the positional trends of multiple first targets, the position of each of the multiple second targets is determined. Since the second target is "MUSH," a set of four characters, the bounding box (e.g., each letter of "MUSH," i.e., M, U, S, H) of each of the multiple second targets is defined as follows: Figure 4A The center point of the small dashed box on the lower right of the "MUSH" target is less than a predetermined distance threshold from the position fitting line (represented by a dashed line with an arrow). This predetermined distance threshold can be 0.1 mm or other values. This is to limit the center point to be near the position fitting line, but not too far away. The center points of the bounding boxes of the three second targets M, U, and S in "MUSH" can be within the bounding boxes of the three first targets. The center point of the bounding box of the first target (the letter H) should be located at or near the center point of the bounding box of the second target (the letter H), which has more bounding boxes than the first target. Ideally, the center point of the bounding box of the second target should be located on the fitted straight line, i.e., the dashed line with the arrow. This maintains the consistency of the text position before and after translation, making it easier for users to read.
[0066] Figure 4B A schematic diagram illustrating yet another example of step 310 of determining a second layout of a plurality of second targets to be included in a second image based on features of a first layout of a plurality of first targets included in a first image, according to at least one embodiment of the present disclosure.
[0067] like Figure 4B The upper part shows the Japanese characters " The center point of the bounding box of each character is taken, and these center points are fitted, which may include non-linear fitting to a curve (indicated by a dashed line with an arrow), which serves as the positional trend of multiple primary targets.
[0068] like Figure 4B As shown in the lower part, based on the positional trends of multiple first targets, the position of each of the multiple second targets is determined. Since the second target is "MUSH," a set of four characters, the distance between the center point of the bounding box of each of the multiple second targets (each letter of "MUSH," i.e., M, U, S, H) and the position fitting line (represented by a dashed line with an arrow) is less than a predetermined distance threshold. The center points of the bounding boxes of the three second targets M, U, and S in "MUSH" can be within the range of the three first targets. The center point of the bounding box of the first target should be at or near the center point of the second target, i.e., the letter H, which has more bounding boxes than the first target. Ideally, the center point of the bounding box should be on the fitted curve, i.e., the dashed line with the arrow. This maintains the consistency of the text position before and after translation, making it easier for users to read.
[0069] Figure 5AA schematic diagram of another example of step 310, according to at least one embodiment of the present disclosure, is shown to determine a second layout of a plurality of second targets to be included in a second image based on features of a first layout of a plurality of first targets included in a first image. Figure 5B A schematic diagram illustrating yet another example of step 310 of determining a second layout of a plurality of second targets to be included in a second image based on features of a first layout of a plurality of first targets included in a first image, according to at least one embodiment of the present disclosure.
[0070] In some embodiments, step 310 of determining a second layout of a plurality of second targets to be included in a second image based on features of a first layout of a plurality of first targets included in a first image may include: determining a rotation angle trend of the plurality of first targets; and determining a rotation angle of each of the plurality of second targets based on the rotation angle trend of the plurality of first targets. Specifically, determining the rotation angle of each of the plurality of second targets based on the rotation angle trend of the plurality of first targets includes making the rotation angle of the second target among the plurality of second targets equal to the rotation angle of the corresponding first target, or the average of the rotation angles of the plurality of first targets, or an angle value whose distance from the rotation angle fitting line is less than a predetermined angle threshold. The predetermined angle threshold may be 0.1 degrees or other values, indicating that the rotation angle can be taken near the rotation angle fitting line, but should not deviate too far from the rotation angle fitting line.
[0071] Here, it should be noted that in one embodiment, if the rotation angle of the second target corresponds to a first target, the selected angle of the first target is used. If there is no corresponding first target (i.e., an extra target, such as the English letter "H"), the selected angle of the extra second target can be the average of the rotation angles of multiple first targets or an angle value whose distance from the rotation angle fitting line is less than a predetermined angle threshold. For example, the first layout contains three first targets, and the rotation angles of each first target are d1, d2, and d3. The average angle is da. If the second layout contains five second targets, then the rotation angle of each second target can be d1, d2, d3, da, and da. Of course, in another embodiment, the rotation angles of all second targets can also be set to the average of the rotation angles of multiple first targets or angle values whose distance from the rotation angle fitting line is less than a predetermined angle threshold. For example, the rotation angles of each second target can be da, da, da, da, and da.
[0072] Determining the rotation angle trend of multiple first targets may include determining the rotation angle of each of the multiple first targets based on the following method.
[0073] like Figure 5AAs shown, images of a standard first target at different rotation angles are compared with the image of the target from among multiple first targets to obtain the similarity between the images of the standard first target at different rotation angles and the image of the target. Here, "standard" first target is used. For example, for this target, the Japanese character "..." ", in its standard Japanese script " "With that first goal, Japanese characters" By comparing the image of "", the standard Japanese character "" can be obtained. Images of different rotation angles (such as) Figure 5A (as shown in the upper part) and the first target " "images (such as Figure 5A The similarity between (as shown in the middle of the image).
[0074] The rotation angle of the first target with the highest similarity is determined as the rotation angle of that first target. For example... Figure 5A As shown, determine the standard Japanese character on the far right. The rotation angle of the character "" is used as the primary objective, the Japanese character "". The rotation angle of ", such as Figure 5A The checkmark icon at the bottom indicates no rotation or a rotation angle of 0. And as... Figure 5A The standard Japanese characters shown by the cross icon at the bottom are " The images of other rotation angles are all related to the first target Japanese character "". Images with a similarity of "" are less similar, meaning they are more different.
[0075] Following the method described above for each first target, the rotation angle of the Japanese character for each first target can be determined. These first targets may have the same or different rotation angles. If the rotation angles of these first targets are basically consistent, the average rotation angle of multiple first targets can be considered as the rotation angle trend of multiple first targets. That is, if multiple first targets are tilted in almost the same direction or not tilted at all, then it is best to generate a second target that is also tilted in that direction or not tilted at all. Figure 5A As shown, if the rotation angle of the Japanese character in each first target is determined to be 0, i.e. no rotation, then the rotation angle of each second target generated by prediction is also 0, i.e. no rotation.
[0076] like Figure 5BAs shown, when the rotation angles of these first targets exhibit a certain trend, such as a radial pattern around a center, or a 10-degree difference between the rotation angles of every two first targets, then a rotation angle fitting line (e.g., a linear line where each target increases by 10 degrees) can be determined as the rotation angle trend of the multiple first targets. It is desirable that the generated second targets also follow this angle trend, such that the rotation angles of the English letters "M", "U", and "S" are the same as the rotation angles of their corresponding three first targets (e.g., each second target increases by 10 degrees), while the rotation angle of the additional English letter "H" is 10 degrees greater than that of the English letter "S". This maintains consistency in the rotation angles of the text before and after translation, facilitating user readability.
[0077] Note that the determination of the position of each of the multiple second targets and the determination of the rotation angle of each second target can be carried out individually or in combination. When carried out in combination, the position of each second target must be determined according to the position trend described above, and the rotation angle of each second target must be determined according to the angle trend described above.
[0078] Thus, by automatically determining the second layout of multiple second targets to be included in the second image based on the objective features of the first layout of multiple first targets included in the first image, the layout style of multiple targets in the generated image can be made as consistent as possible with the original layout style, and manual adjustment can be omitted, thereby achieving efficient and optimized layout of multiple targets.
[0079] In some embodiments, the total bounding box of the multiple second targets needs to be limited to or not significantly different from the total bounding box of the multiple first targets, in order to prevent the newly generated multiple second targets from covering surrounding areas that should not have been covered.
[0080] Figure 6 A schematic diagram is shown illustrating, according to at least one embodiment of the present disclosure, that the total bounding box of a plurality of second targets is limited to or is not significantly different from the total bounding box of a plurality of first targets.
[0081] like Figure 6 As shown in the upper part, the extra second target (see the small dashed box) exceeds the range of the total bounding box of the multiple first targets (see the large dashed box), causing the range of the total bounding box of the multiple second targets to exceed the range of the total bounding box of the multiple first targets.
[0082] The bounding box of at least one of multiple secondary targets can be reduced (e.g.) Figure 6(As shown in the lower part), so that the difference or ratio between the total bounding box range of the multiple second targets and the total bounding box range including the multiple first targets is not greater than a predetermined threshold. Here, the bounding box reduction can be proportional or non-proportional.
[0083] Alternatively, the bounding boxes of at least one of the multiple second targets can be overlapped (or overlapped more) so that the difference or ratio between the total bounding box range of the multiple second targets and the total bounding box range including the multiple first targets is not greater than a predetermined threshold.
[0084] Alternatively, the bounding boxes of some second targets can be reduced, some second targets can be overlapped, or other methods can be used to ensure that the difference or ratio between the total bounding box range of multiple second targets and the total bounding box range including multiple first targets is not greater than a predetermined threshold. As described above, the predetermined threshold can be 1.05 times, but of course, this value is not limited to this, as long as it avoids the newly generated multiple second targets from covering the surrounding areas that should not have been covered in the first place.
[0085] However, sometimes, if the second targets overlap too much, the generated second targets may not be displayed correctly, display errors, or be severely occluded, resulting in image quality problems and causing generation failure or errors.
[0086] Figure 7 This diagram illustrates a situation where excessive overlap between second targets causes the generated second targets to fail to display correctly or display incorrectly. For example... Figure 7 As shown, due to excessive overlap between the bounding boxes of the multiple targets to be generated or other reasons, the three targets (M, U, and S) in the generated image cannot be displayed normally, and only one target (i.e., the letter "S") is fully displayed. This results in image quality problems, leading to generation failure or errors.
[0087] In some embodiments, step 310 of determining the second layout of a plurality of second targets to be included in a second image based on the features of the first layout of a plurality of first targets included in a first image may include: determining the position and size of the bounding box of each of the plurality of second targets, such that the union of the bounding boxes of the plurality of second targets is within the range of the total bounding box of the plurality of second targets and has the largest area, and the overlap area between the bounding boxes of the plurality of second targets is the smallest.
[0088] Here, the "union" of bounding boxes refers to the smallest bounding shape that is typically used to calculate the result of merging two or more bounding boxes.
[0089] Figure 8A schematic diagram is shown according to at least one embodiment of the present disclosure, in which the union of the bounding boxes of a plurality of second targets is within the total bounding box of the plurality of second targets and has the largest area, and the area of overlap between the bounding boxes of the plurality of second targets is minimized.
[0090] In some embodiments, determining the position and size of the bounding box of each of the plurality of second targets, such that the union of the bounding boxes of the plurality of second targets is within the range of the total bounding box of the plurality of second targets and has the largest area, and the overlap area between the bounding boxes of the plurality of second targets is the smallest, may include the following steps.
[0091] Determine the initial position and initial size of the bounding box for each of the multiple second targets. This initial position and initial size may be the layout of the multiple second targets as described above, determined based on the positional and / or angular trends of the multiple first targets, or other initial settings.
[0092] Then, the position and / or size of the bounding boxes of at least one of the multiple second targets are gradually moved and / or changed, so that the union of the bounding boxes of the multiple second targets is within the total bounding box of the multiple second targets with the largest area, and the overlap between the bounding boxes of the multiple second targets is minimized. Here, it should be noted that when both the first target and the second target are characters of language, in order to maintain the language habits of the second targets and make them easy for humans to read, the relative positions and / or sizes of the bounding boxes of the second targets in the original second layout can be kept unchanged as much as possible.
[0093] like Figure 8 The left side shows, for example, in the initial layout before adjusting the layout of the second targets, the overlap between the three second targets ("M", "U", "S") is quite large, while... Figure 8 As shown on the right, after adjusting the layout of the second targets in a way that maximizes the area of the union of the bounding boxes ("M", "U", "S") of the multiple second targets within the total bounding box of the multiple second targets, and minimizes the area of overlap between the bounding boxes of the multiple second targets, the overlap area between the three second targets becomes smaller, making them less likely to occlude each other. Moreover, the total area occupied by the three second targets becomes larger, making them more clearly visible. This ensures that each second target achieves an optimal level of clarity and is less likely to occlude each other, making it easier to ensure that the generated second targets will not fail or err.
[0094] Note that in Figure 8 The target shown is an English letter, but the types of targets are not limited to this; they can also be people, animals, still life, etc.
[0095] Figure 9 A schematic diagram is shown illustrating the use of a genetic algorithm according to at least one embodiment of the present disclosure to maximize the area of the union of the bounding boxes of a plurality of second targets within the total bounding box of the plurality of second targets, and minimize the area of overlap between the bounding boxes of the plurality of second targets.
[0096] like Figure 9As shown, determining the position and size of the bounding box for each of multiple secondary objectives, such that the union of the bounding boxes of the multiple secondary objectives lies within the total bounding box of the multiple secondary objectives with the largest area, and the overlap between the bounding boxes of the multiple secondary objectives is minimized, can be achieved using a Genetic Algorithm (GA). A Genetic Algorithm is an optimization algorithm that simulates the Darwinian evolutionary process of gene sequences, mainly including three operations: selection, crossover, and mutation. These operations simulate the processes of natural selection and genetic mutation, helping the algorithm gradually approach the optimal solution in the search space. The selection operation selects superior individuals as parents based on their fitness value (i.e., the objective function value of the problem). Tournament selection can be used for this selection. Individuals with higher fitness values have a greater probability of being selected. The crossover operation pairs the selected parent individuals and generates new offspring individuals by exchanging some gene segments. This process simulates gene recombination in biological evolution. The mutation operation performs small-probability random changes on the offspring individuals to increase the diversity of the population. The basic process of a genetic algorithm includes population initialization, decoding and fitness calculation, selection, crossover, mutation, and termination condition determination. Population initialization involves randomly generating a certain number of individuals as initial solutions. Each individual consists of a set of gene sequences (or chromosomes), which determine the individual's characteristics. Decoding and fitness calculation involves decoding each individual in the population, transforming it into a solution in the problem space, and calculating its fitness value. The fitness value is used to evaluate the quality of the solution. Selection involves selecting superior individuals as parents based on their fitness values to generate the next generation. Crossover involves performing a crossover operation on the selected parent individuals to generate new offspring individuals. Mutation involves performing a mutation operation on the offspring individuals to increase the diversity of the population. Termination condition determination involves determining whether the algorithm has reached a preset stopping condition, such as reaching the maximum number of iterations or finding a solution that meets the requirements. If the condition is met, the algorithm terminates and outputs the optimal solution; otherwise, it continues iterating. Detailed steps are as follows: Population initialization: The bounding boxes in the initial layout are considered as gene sequences. Multiple gene sequences constitute an individual (layout), and multiple individuals constitute the population. Prior knowledge-guided mutation and crossover operations: Prior knowledge is used to regulate random operations (translation and scaling) on individuals during genetic evolution. Specifically, translation: the new center point of each bounding box is calculated by offsetting the original center point. Scaling: the new height and width of each bounding box are calculated by scaling the original height and width by a certain proportion. The selection method uses a traditional tournament selection method.The choice of a new fitness function—(sum of the areas of all bounding boxes in the layout) - (sum of the overlapping areas between any two bounding boxes in the layout) or (sum of the areas of all bounding boxes in the layout) / (sum of the overlapping areas between any two bounding boxes in the layout)—allows for a large coverage area of the bounding boxes and a small overlap area between them. Termination condition: The algorithm terminates when the preset maximum number of iterations is reached or when the fitness value shows no significant improvement for several consecutive generations.
[0097] In this embodiment, the bounding box of each of the multiple second objectives can be used as a gene sequence in the genetic algorithm, the layout of the bounding boxes of the multiple second objectives can be used as an individual, and the ratio or difference between the area of the union of the bounding boxes of the multiple second objectives and the area of the overlap between the bounding boxes of the multiple second objectives can be used as a fitness value. Crossover and mutation can be used to cross and change the position and / or size of the bounding boxes of one or more of the multiple second objectives.
[0098] Thus, the application of the genetic algorithm in this embodiment may specifically include: using the bounding box of each of the multiple second objectives as a gene sequence in the genetic algorithm, and the layout of the bounding boxes of the multiple second objectives as an individual; in each iteration of the genetic algorithm, performing random operations on the individuals according to prior knowledge specifications, including translation and scaling, wherein translation includes translating the position of the bounding box in the individual, and scaling includes scaling the size of the bounding box in the individual proportionally, wherein the prior knowledge specifications include keeping the relative positional relationship between the positions of the bounding boxes of the multiple second objectives unchanged and the relative size relationship between the sizes of the bounding boxes of the multiple second objectives unchanged; selecting individuals with higher fitness, wherein fitness is the ratio or difference between the area of the union of the bounding boxes of the multiple second objectives and the area of the overlap between the bounding boxes of the multiple second objectives; randomly changing the position and / or size of the bounding boxes in the selected individuals with a predetermined probability to generate a new set of individuals in the next iteration; stopping the iteration in response to the number of iterations reaching a predetermined threshold or in response to the fitness no longer increasing.
[0099] Thus, genetic algorithms can be used to achieve, more quickly and reasonably, the union of the bounding boxes of multiple second objectives within the total bounding box of the multiple second objectives with the largest area and the smallest overlap between the bounding boxes of the multiple second objectives.
[0100] Of course, there are many other algorithms that maximize the area of the union of the bounding boxes of multiple second targets within the total bounding box of the multiple second targets, and minimize the area of overlap between the bounding boxes of the multiple second targets, which will not be listed here.
[0101] Next, after determining the second layout for multiple secondary objectives, such as Figure 3As shown, in step 320, a second image including multiple second targets is generated based on the second layout and the first image.
[0102] In some embodiments, step 320 of generating a second image including multiple second targets based on a second layout and a first image may include: cropping multiple first targets from the first image and filling the background of the cropped portions of the multiple first targets to obtain a first image with the first targets removed; and combining the multiple second targets according to the second layout on the first image with the first targets removed to generate a second image including multiple second targets. The multiple second targets can be derived from multiple first targets using human or artificial intelligence technology, such as human translation or automatic translation using artificial intelligence technology. Of course, other methods may also be used to generate a second image including multiple second targets, which will not be listed here.
[0103] Thus, taking the example of translating Japanese into English, the Japanese text can be removed from the original Japanese image, and the translated English letters can be combined onto the original Japanese image with the Japanese text removed according to a determined second layout to generate a second image including the translated English text. This process can be repeated continuously, efficiently, quickly, and accurately converting an entire Japanese picture book into an English picture book without affecting the reading experience by covering unnecessary areas due to an increase in the number of English letters. As described earlier, in the experiment, manually adjusting 1000 images took approximately 2.7 hours, while generating 1000 images of English text using the solution according to at least one embodiment of this disclosure took approximately 0.5 hours, increasing efficiency by approximately 5.4 times, which greatly improves efficiency.
[0104] According to at least one embodiment of the present disclosure, an apparatus for generating an image including a plurality of targets may also be provided, including components that perform the steps of a method for generating an image including a plurality of targets according to at least one embodiment of the present disclosure. A description of the steps of a method for generating an image including a plurality of targets according to at least one embodiment of the present disclosure can be found above.
[0105] This disclosure may also include non-transitory computer-readable storage media. Instructions, such as computer instructions, are stored on the non-transitory computer-readable storage medium. When the computer instructions are executed by a processor, the various methods described above can be performed. Non-transitory computer-readable storage media include, but are not limited to, random access memory (RAM), read-only memory (ROM), flash memory, EPROM memory, EEPROM memory, registers, computer storage media (e.g., hard disks, floppy disks, solid-state drives, removable disks, CD-ROMs, DVD-ROMs, Blu-ray discs, etc.). For example, the non-transitory computer-readable storage medium can be connected to a computing device such as a computer, and then, when the computing device executes the computer instructions stored on the computer-readable storage medium, the various methods described above can be performed.
[0106] This disclosure may also include a computer program product capable of performing the methods, steps, and operations given herein. For example, such a computer program product may be a computer software package, computer code instructions, or a computer-readable tangible medium having computer instructions tangibly stored (and / or encoded) thereon, which can be executed by a processor to perform the operations described herein. The computer program product may include packaging materials.
[0107] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The term “such as / for example” as used herein refers to the phrase “such as / for example but not limited to,” and is used interchangeably with it.
[0108] The flowcharts and method descriptions in this disclosure are merely illustrative examples and are not intended to require or imply that the steps of the various embodiments must be performed in the given order. As those skilled in the art will recognize, the steps in the above embodiments can be performed in any order. Words such as "then," "next," etc., are not intended to limit the order of the steps; these words are only used to guide the reader through the description of these methods. Furthermore, any reference to a singular element, such as the use of the articles "a," "one," or "the," is not to be construed as limiting that element to the singular.
[0109] Furthermore, the steps and apparatus in the various embodiments herein are not limited to any one embodiment. In fact, new embodiments can be conceived by combining relevant steps and apparatus in the various embodiments herein based on the concepts of this disclosure, and these new embodiments are also included within the scope of this disclosure.
[0110] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit at least one embodiment of the present disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A method for generating an image comprising multiple targets, comprising: A second layout of multiple second targets to be included in a second image is determined based on the features of a first layout of multiple first targets included in a first image, wherein the number of the multiple second targets is greater than the number of the multiple first targets, including that the difference or ratio between the range of the total bounding box of the multiple second targets and the range of the total bounding box including the multiple first targets is not greater than a predetermined threshold; and A second image including the plurality of second targets is generated based on the second layout and the first image. The step of determining the second layout of the multiple second targets to be included in the second image based on the features of the first layout of the multiple first targets included in the first image includes: Determine at least the positional trend and / or rotational angle trend of the plurality of first targets; Based on the position trends and / or rotation angle trends of the plurality of first targets, determine the position and / or rotation angle of each of the plurality of second targets.
2. The method according to claim 1, wherein, Determining at least the positional trend and / or rotation angle trend of the plurality of first targets includes: Determine the position fitting line of the center point of the bounding box of each of the plurality of first targets, as the position trend of the plurality of first targets. The determination of the position and / or rotation angle of each of the plurality of second targets based on the position trends and / or rotation angle trends of the plurality of first targets includes: This ensures that the distance between the center point of the bounding box of each of the plurality of second targets and the position fitting line is less than a predetermined distance threshold.
3. The method according to claim 1, wherein, Determining at least the positional trend and / or rotation angle trend of the plurality of first targets includes: The rotation angle of each of the plurality of first targets is determined based on the following method: The images of the standard first target at different rotation angles of one of the plurality of first targets are compared with the image of the first target to obtain the similarity between the images of the standard first target at different rotation angles and the image of the first target. The rotation angle of the standard first target with the highest similarity is determined as the rotation angle of the first target; The average value or a rotation angle fitting line of the plurality of first targets is determined as the rotation angle trend of the plurality of first targets. The step of determining the position and / or rotation angle of each of the plurality of second targets based on the position trends and / or rotation angle trends of the plurality of first targets includes: The rotation angle of the second target among the plurality of second targets is such that the rotation angle of the corresponding first target, or the average value, or the angle value whose distance from the rotation angle fitting line is less than a predetermined angle threshold.
4. The method according to claim 1, wherein, The step of determining the second layout of multiple second targets to be included in the second image based on the features of the first layout of multiple first targets included in the first image includes: Shrink the bounding box of at least one of the plurality of second targets, and / or overlap the bounding boxes of at least one of the plurality of second targets, such that the difference or ratio between the total bounding box range of the plurality of second targets and the total bounding box range including the plurality of first targets is not greater than a predetermined threshold.
5. The method according to claim 1, wherein, The step of determining the second layout of multiple second targets to be included in the second image based on the features of the first layout of multiple first targets included in the first image includes: Determine the position and size of the bounding box of each of the plurality of second targets, such that the union of the bounding boxes of the plurality of second targets is within the total bounding box of the plurality of second targets and has the largest area, and the overlap area between the bounding boxes of the plurality of second targets is the smallest.
6. The method according to claim 5, wherein, Determining the position and size of the bounding box of each of the plurality of second targets, such that the union of the bounding boxes of the plurality of second targets is within the total bounding box of the plurality of second targets and has the largest area, and the overlap area between the bounding boxes of the plurality of second targets is the smallest, includes: Determine the initial position and initial size of the bounding box of each of the plurality of second targets; Move the position of the bounding box of at least one of the plurality of second targets and / or change the size of the bounding box such that the union of the bounding boxes of the plurality of second targets is within the total bounding box of the plurality of second targets and has the largest area, and the overlap area between the bounding boxes of the plurality of second targets is the smallest.
7. The method according to claim 5, wherein, Determine the position and size of the bounding box of each of the plurality of second targets, such that the union of the bounding boxes of the plurality of second targets is within the total bounding box of the plurality of second targets with the largest area, and the overlap area between the bounding boxes of the plurality of second targets is minimized. This is achieved through a genetic algorithm, including: The bounding box of each of the plurality of second objectives is used as the gene sequence in the genetic algorithm, and the layout of the bounding boxes of the plurality of second objectives is used as an individual; In each iteration of the genetic algorithm, the individual is subjected to random operations according to the prior knowledge specification. The operations include translation and scaling. The translation includes translating the position of the bounding box in the individual. The scaling includes scaling the size of the bounding box in the individual proportionally. The prior knowledge specification includes keeping the relative positional relationship between the bounding boxes of the multiple second targets unchanged and the relative size relationship between the bounding boxes of the multiple second targets unchanged. Select individuals with higher fitness, wherein fitness is the ratio or difference between the area of the union of the bounding boxes of the plurality of second targets and the area of the overlap between the bounding boxes of the plurality of second targets. The position and / or size of the bounding box in the selected individuals are randomly changed with a predetermined probability to generate a new set of individuals in the next iteration; The iteration stops when the number of iterations reaches a predetermined threshold or when the fitness no longer increases.
8. The method according to claim 1, wherein, The plurality of first targets are multiple characters of a first language, and the plurality of second targets are multiple characters of a second language, wherein the multiple characters of the second language are translations of the multiple characters of the first language, and the arrangement of the plurality of first targets in the first layout and the arrangement of the plurality of second targets in the second layout respectively conform to the language habits of the first language and the language habits of the second language.
9. The method according to claim 1, wherein, The step of generating a second image including the plurality of second targets based on the second layout and the first image includes: The plurality of first targets are cropped from the first image and the portion from which the plurality of first targets are cropped is filled with background to obtain a first image with the first targets removed; The plurality of second targets are combined on the first image, from which the first target has been removed, according to the second layout to generate a second image including the plurality of second targets.
10. An apparatus for generating an image comprising a plurality of targets, comprising: Memory, which stores computer instructions; At least one processor is configured to execute the computer instructions in the memory to perform the method according to any one of claims 1-9.