Learning device, inference device, and image processing device
The learning and inference devices enhance character visibility evaluation by considering positional relationship, attributes, color, mode, and medium, addressing the limitations of existing methods to improve display efficiency.
Patent Information
- Application Number
- JP2022045329
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-22
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2042-03-22
Smart Images

Figure 0007752555000001 
Figure 0007752555000002 
Figure 0007752555000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a learning device, an inference device, and an image processing device including an inference device. [Background technology]
[0002] As an image processing device for arranging many figures and characters in a limited area, Japanese Patent Application Laid-Open No. 2011-33705 (Patent Document 1) describes a method for efficiently displaying multiple display objects by introducing a judgment figure to determine whether the display objects overlap each other when displaying multiple characters or the like that are to be displayed overlapping each other.
[0003] In the image processing device of Patent Document 1, when multiple display objects are displayed overlapping each other, the judgment figures, which are created to have an area smaller than the circumscribing figures of each display object, are arranged so that they do not overlap, thereby enabling the multiple display objects to be displayed efficiently. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2011-33705 Summary of the Invention [Problem to be solved by the invention]
[0005] The image processing device of Patent Document 1 adjusts the relative positional relationship of the display areas so that the determination figures determined corresponding to the display areas of multiple characters, etc. do not overlap, thereby formulating a layout that allows each of the multiple display objects displayed on top of each other to be identified and visible.
[0006] However, from the viewpoint of ensuring human visibility, factors other than the relative positional relationship of the display areas may also have an impact. In other words, a determination that focuses only on the relative positional relationship of the display areas, as in Patent Document 1, may inappropriately evaluate whether multiple overlapping characters are distinguishable and legible. As a result, there is a concern that even if the determination figures are arranged so that they do not overlap, visibility may be reduced depending on the display mode, or that display efficiency may be reduced as a result of a layout in which the determination figures do not overlap.
[0007] The present disclosure has been made to solve such problems, and the purpose of the present disclosure is to provide a learning device, an inference device, and an image processing device that are capable of appropriately evaluating the visibility of multiple characters that are displayed in an overlapping manner according to image data. [Means for solving the problem]
[0008] In one aspect of the present disclosure, a learning device is provided. The learning device includes a data acquisition unit and a model generation unit. The data acquisition unit acquires learning data for each of a plurality of characters displayed in an overlapping manner by combining a first input, a second input, a third input, and a fourth input from image data indicating the display content of the characters. The model generation unit generates a trained model for inferring whether each character is visible and distinguishable from other characters displayed in an overlapping manner, from the combination of the first input, the second input, the third input, and the fourth input included in the learning data. The first input includes information related to the relative positional relationship of the plurality of characters and information related to the display attributes of the characters. The second input includes information related to the display color of the characters. The third input includes information related to the display mode of the characters. The fourth input includes information related to the display medium of the characters.
[0009] In another aspect of the present disclosure, an inference device is provided. The inference device includes a data acquisition unit and an inference unit. The data acquisition unit acquires model input data for each of a plurality of characters displayed in an overlapping manner by combining a first input, a second input, a third input, and a fourth input from image data showing the display content of the characters. The inference unit inputs the model input data to a trained model for inferring whether each character is distinguishable from other characters displayed in an overlapping manner, based on the first input, the second input, the third input, and the fourth input, and outputs an inference result indicating whether each of the plurality of characters displayed in accordance with the image data is distinguishable and visible. The first input includes information related to the relative positional relationship of the plurality of characters and information related to the display attributes of the characters. The second input includes information related to the display color of the characters. The third input includes information related to the display mode of the characters. The fourth input includes information related to the display medium of the characters.
[0010] In another aspect of the present disclosure, an image processing device is provided. The image processing device includes the inference device described above, an input unit, a display data storage unit, and a display control unit. The input unit specifies multiple characters to be displayed in an overlapping manner as display targets. The display data storage unit stores data including a first input, a second input, a third input, and a fourth input when displaying the display targets. The display control unit generates information on whether each of the multiple characters is distinguishable and visible from an output of the inference device obtained by inputting the first input, the second input, the third input, and the fourth input for each of the multiple characters to the inference device. [Effects of the Invention]
[0011] According to the present disclosure, it is possible to provide a learning device, an inference device, and an image processing device that can appropriately evaluate the visibility of multiple characters that are displayed in an overlapping manner according to image data by taking into account information on the relative positional relationship of multiple characters as well as information on the display attributes, display color, display mode, and display medium of the characters. [Brief explanation of the drawings]
[0012] [Figure 1] 1 is a block diagram of a learning device according to a first embodiment. [Figure 2] FIG. 10 is a conceptual diagram showing an example of a display mode of a plurality of characters displayed in an overlapping manner. [Figure 3] 10 is a diagram illustrating an example of the configuration of input data acquired by a data acquisition unit. [Figure 4] 4 is a flowchart illustrating a learning method in the learning device according to the first embodiment. [Figure 5] FIG. 10 is a schematic block diagram of a learning device according to a modification of the first embodiment. [Figure 6] FIG. 10 is a conceptual diagram for explaining a determination figure. [Figure 7] FIG. 10 is a conceptual diagram illustrating an example of the relationship between a determination figure and a circumscribing figure. [Figure 8] FIG. 10 is a block diagram illustrating an example of the configuration of an inference device and an image processing device according to a second embodiment. [Figure 9] 10 is a flowchart illustrating a first example of an inference process performed by the inference device. [Figure 10] 10 is a flowchart illustrating a second example of inference processing by the inference device. [Figure 11] FIG. 10 is a block diagram illustrating an example of the configuration of an inference device and an image processing device according to a modified example of the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0013] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the following, the same or corresponding parts in the drawings will be denoted by the same reference numerals, and their description will not be repeated in principle.
[0014] Embodiment 1 In the first embodiment, a learning model applied to an image processing device according to the present embodiment will be described.
[0015] FIG. 1 shows a block diagram of a learning device 100 according to the first embodiment. Learning device 100 receives as input electronic data (hereinafter referred to as "image data") for image output used in an image processing device that displays characters and the like on a display medium (not shown) such as a display or printer, and generates a trained model for inferring character strings and groups of figures that can ensure visibility even when there is overlap between displayed characters. The image data can be, for example, design data such as CAD (Computer Aided Design), or electronic data for issuing drawing instructions to a printer, display, etc. that has been converted from electronic data, etc., via the Internet.
[0016] Learning device 100 is configured, for example, by a personal computer (not shown) having a central processing unit (CPU), memory, an input / output interface, and a display, which executes a pre-installed program. Alternatively, the functions of learning device 100 may be realized on the cloud.
[0017] 1, a learning device 100 according to the first embodiment includes a data acquiring unit 110 and a model generating unit 120. The data acquiring unit 110 receives inputs 1 to 4 extracted from image data that is data to be output to the display medium, and generates learning data xi.
[0018] The image data can be digital data collected by the image processing device by reading data stored in an auxiliary storage device such as a hard disk of the image processing device or by downloading via the Internet. The digital data includes various information for specifying the display mode, including data indicating the contents of inputs 1 to 4 described below.
[0019] For example, characters whose display areas overlap are extracted from the entire image data, and then data portions corresponding to inputs 1 to 4 (described later) are extracted from the image data for each of the extracted characters and input to the data acquisition unit 110. Note that the overlapping of display areas here may simply refer to overlaps between the fill area of a target character and the fill area of another character, or may refer to overlaps between the outline area of a target character and the outline area of another character. Based on this combination of target characters, inputs 1, 2, 3, and 4 are extracted for each character.
[0020] In this embodiment, the i-th (i: natural number) learning data xi is generated by combining inputs 1 to 4 in correspondence with each of the multiple characters displayed in an overlapping manner.
[0021] 2 is a conceptual diagram showing an example of a display mode of a plurality of characters to be learned, which are displayed overlapping each other, in which two-dimensional character display is shown as a display example.
[0022] 2, for two characters that are displayed overlapping each other, the display area 10 for the character data of "A" overlaps with the display area 20 for "B." The coordinate Pa(xa, ya) indicates the center point of the character "A" in the display area 10 on the display medium (two-dimensional), and similarly, the coordinate Pb(xb, yb) indicates the center point of the character "B" in the display area 20.
[0023] 2, the display areas 10 and 20 are drawn in white to make the positions of the coordinates Pa and Pb of the center points easier to understand, but in reality, the characters "A" and "B" are displayed in the display areas 10 and 20, respectively, filled in with a specified character color. The above-mentioned learning data xi are generated individually for the character "A" having the display area 10 and the character "B" having the display area 20.
[0024] FIG. 3 shows a diagram illustrating an example of the configuration of inputs 1 to 4 that make up the learning data.
[0025] 3, input 1 includes information indicating the basic attributes of each character to be displayed in an overlapping manner. Input 1 includes information indicating the relative positional relationship between the character and other characters to be displayed in an overlapping manner, and information indicating the display attributes of each character, such as data indicating the font type and character size.
[0026] The relative coordinates are calculated as data quantitatively indicating the relative positional relationship between characters displayed on top of each other, for example, as the difference (xb-xa, yb-ya) between the coordinates Pa and Pb in Figure 2. The magnitude of the position vector of the relative coordinates is √((xb-xa) 2 +yb-ya) 2 ) indicates that the distance between characters is small. When characters are closer to each other or when the overlap between characters decreases, visibility decreases. In the display example of Figure 2, for each of the characters "A" and "B," the relative coordinates (xb-xa, yb-ya) are acquired as part of input 1.
[0027] Font type is indicated by specifying "Mincho" or "Gothic" etc. Character size is indicated by the font size. Differences in character shape due to font type affect visibility. Specifically, when bold characters overlap, visibility tends to decrease. Also, since visibility generally decreases as the font size becomes smaller, even when the degree of overlap is small, visibility tends to decrease if the font size is small.
[0028] By setting Input 1 in this way, in addition to the relative positional relationship of multiple characters displayed on top of each other, the font type and size of the characters can also be included in the training data as data that affect visibility.
[0029] Input 2 includes information related to display colors. Specifically, input 2 includes data indicating the display color of each character and data indicating the background color of the character. When the display colors of overlapping characters, or the display color and background color, are similar colors, such as red and orange, visibility tends to decrease. Thus, the display color and background color are also parameters that affect visibility, and are therefore included in the training data. For example, each of the display color and background color can be represented by a quantitative value (data) indicating coordinates in a color space such as RGB (Red Green Blue). In this case, for each overlapping character, the R value, G value, and B value of character color 1, which is the display color of the character, the background color of the character, and character color 2, which is the display color of the other character that is displayed overlapping, are input to the model generation unit 120 as input 2.
[0030] Any known color space can be used, such as CMY (Cyan Magenta Yellow), YIQ, or HSV (Hue Saturation Value). It is also possible to add data values indicating luminance or brightness to the input 2 items. By setting input 2 in this way, it is possible to include in the training data information indicating characteristics that affect the visibility of overlaid characters from the perspective of display color (hue).
[0031] Input 3 is information indicating the display format of the character string to be learned. For example, Input 3 includes information indicating whether the display format is a moving image or a still image, and information indicating whether the display format is a two-dimensional display (2D) or a three-dimensional display (3D).
[0032] Generally, the visibility of overlaid character strings is lower in moving images than in still images. Also, character strings displayed in three dimensions have lower visibility than character strings displayed in two dimensions. By setting input 3 in this way, it is possible to learn and input data that indicates characteristics that affect the visibility of overlaid character strings from the perspective of display format.
[0033] Alternatively, the input 3 may include a use code indicating the display use. For example, the use code may have different values predefined for each use, such as to indicate a pre-specified use such as displaying on a map or displaying a blueprint.
[0034] Input 4 includes information indicating the type or characteristics of the display medium. For example, input 4 includes data indicating whether the display medium is a display or a printed matter (paper, plastic plate, metal plate, etc.), and data indicating the display area (size) of the display or printed matter.
[0035] Furthermore, the input 4 may further include data indicating the resolution for a display and data indicating the material of the printed surface for a printed matter. The data indicating the material preferably includes data indicating the surface properties (such as glossiness).
[0036] When displayed on a display, the higher the resolution, the better the visibility, and when printed on paper, the size and surface texture of the printed surface affect visibility. By setting input 4 in this way, it is possible to learn and input data that indicates the characteristics that affect the visibility of overlaid characters depending on the type of display medium.
[0037] Referring again to FIG. 1, as described above, the data acquisition unit 110 generates learning data xi (i: natural number) for each of the multiple characters displayed in an overlapping manner by combining the above-mentioned inputs 1 to 4 in accordance with a predetermined format.
[0038] The model generation unit 120 generates a trained model using the training data xi from the data acquisition unit 110 as input. The trained model storage unit 300 stores the trained model generated by the model generation unit 120.
[0039] The model generation unit 120 can generate a trained model that provides an output 1 for training data xi (input 1 to input 4) using a known training algorithm. The output 1 includes a determination result as to whether the multiple characters displayed overlapping each other, indicated by input 1 to input 4, are distinguishable and visible (hereinafter simply referred to as "legible") or distinguishable and invisible (hereinafter simply referred to as "unreadable").
[0040] When AI (Artificial Intelligence) is used to determine the visibility of overlapping character strings, that is, whether they are human-readable or not, the trained model is configured as a model for classifying (clustering) combinations of inputs 1 to 4 in which characters can be ensured to be visible relative to other overlapping characters (when legible) and combinations of inputs 1 to 4 in which characters cannot be ensured to be visible (when illegible).
[0041] As described above, the learning algorithm used by the model generation unit 120 can be a known algorithm such as supervised learning, unsupervised learning, or reinforcement learning. For example, the model generation unit 120 can learn the output 1 by so-called unsupervised learning in accordance with a clustering method using the k-means method. Note that unsupervised learning refers to a method of providing a learning device with learning data that does not include results (labels) and learning features of the learning data.
[0042] The k-means method is a hierarchical clustering algorithm that classifies a given number of clusters into k (k: natural number) using the cluster mean. Specifically, the k-means method can be executed for each of the learning data xi of the overlapping characters mentioned above as follows:
[0043] First, N pieces of input learning data xi (i: 1 to N) are randomly assigned to clusters. For example, there can be two clusters (k=2): one in which the character is legible when compared to other characters displayed over it (hereinafter also referred to as a "legible cluster"), and one in which the character is illegible (hereinafter also referred to as an "indecipherable cluster").
[0044] Next, under the current allocation status of the N pieces of training data xi, the center Vj (j: 1≦j≦k) of each of the k clusters is calculated. Vj corresponds to a combination of the average values of each data item of the training data xi assigned to that cluster. Furthermore, for each piece of training data xi, the distance from the center Vj of each of the k clusters is calculated, and the training data xi is reassigned to the cluster closest to the center value.
[0045] By repeatedly changing the allocation of each learning data xi in this way, when the cluster allocation of all learning data no longer changes or the rate of change in allocation becomes lower than a predetermined judgment value, it is determined that learning has converged and a learned model is created.
[0046] In unsupervised learning, image output data of character strings evaluated as readable is input to the data acquisition unit 110, and a trained model is created by the model generation unit 120. Furthermore, supervised learning can be performed by preparing image data of character strings that are intentionally evaluated as unreadable in addition to image data of character strings evaluated as readable and inputting the image data to the data acquisition unit 110.
[0047] Furthermore, in the trained model, a boundary W between legible and unreadable, which corresponds to the boundary between the two clusters described above, can be further determined. The boundary W is also defined as a combination of the data values (boundary values) of each item of input 1 to input 4 described above, which are components of the training data xi.
[0048] Therefore, the output 1 of the trained model can include, in addition to the determination result of whether the data is readable or unreadable as described above, the amount of correction of inputs 1 to 4 to make the training data xi readable in accordance with its relative relationship with the boundary W described above.
[0049] Fig. 4 shows a flowchart illustrating a learning method in the learning device according to embodiment 1. For example, the control process in Fig. 4 can be executed by a computer on which a program for executing the functions of learning device 100 is installed.
[0050] 4, learning device 100 acquires data relating to inputs 1 to 4 for each of a plurality of characters displayed overlappingly and included in the display object in step (hereinafter simply referred to as "S") 10. The processing of S10 is realized by extracting data for the items included in inputs 1 to 4 from the image data of the display object.
[0051] In S12, the learning device 100 generates learning data xi for each character to be displayed in an overlapping manner from the data of inputs 1 to 4 acquired in S10. Note that in S10, inputs 1 to 4 may be acquired in parallel for each character, or data for multiple characters of a string to be displayed in an overlapping manner for at least one of inputs 1 to 4 may be acquired in parallel with their corresponding relationships with each character indicated. In this case, in S12, learning data xi for each character is generated, including processing for classifying the data for multiple characters acquired in parallel by display object.
[0052] In S14, the learning device 100 generates a trained model for obtaining output 1 from inputs 1 to 4 described above through a learning process in the model generation unit 120 using the training data generated in S12. Then, in S16, the learning device 100 stores the trained model. In FIG. 1, the trained model storage unit 300, which is shown as the storage destination for the trained model, may be configured as a storage device external to a computer (not shown) that realizes the functions of the learning device 100, or may be a storage device internal to the computer. The trained model can also be deployed to a storage device of another computer via any storage medium or a communication network.
[0053] In this way, the learning device according to embodiment 1 can generate a trained model that can accurately determine whether multiple overlapping characters are legible, taking into consideration not only the relative positional relationship between the overlapping characters, but also character attributes (input 1) such as character shape and size that affect visibility when overlapping characters, display color (input 2) that affects contrast, display format (input 3), and the characteristics of the display medium (input 4), which affect the relative positional relationship between the overlapping characters.
[0054] A variation of the first embodiment. It is possible to further combine the learning related to the variables of the judgment graphic described in Patent Document 1 with the first embodiment.
[0055] FIG. 5 is a schematic block diagram of a learning device according to a modification of the first embodiment. 5, learning device 101 according to the first embodiment further includes a judgment graphic database 130 in addition to data acquiring unit 110 and model generating unit 120 similar to those of learning device 100 shown in FIG.
[0056] The judgment figure is similar to that disclosed in Patent Document 1. Fig. 6 shows a conceptual diagram for explaining the judgment figure.
[0057] FIG. 6 shows an external circumscribing figure 302a and a determination figure 303a for the display character 301a, and an external circumscribing figure 302b and a determination figure 303b for the display character 301b, in a display target where the display characters 301a (the Ming-style 'A') and 301b (the Ming-style 'B') are superposed and displayed.
[0058] As described in Patent Document 1, the external circumscribing figure 302a is defined as a rectangular shape that circumscribes the display character 301a. The determination figure 303a is defined as a figure that has a smaller area than the external circumscribing figure 302a and is included inside the external circumscribing figure 302a. The relationship among the display character 301b, the external circumscribing figure 302b, and the determination figure 303b is the same as the relationship among the display character 301a, the external circumscribing figure 302a, and the determination figure 303a.
[0059] FIG. 7 shows a conceptual diagram for explaining an example of the relationship between the determination figure and the external circumscribing figure. In FIG. 7, the external circumscribing figure 302 is a general term for the external circumscribing figures 302a and 302b in FIG. 6, and the determination figure 303 is a general term for the determination figures 303a and 303b in FIG. 6.
[0060] The rectangular external circumscribing figure 302 has a long side wz (xy) and a short side wx (zy). In contrast, the determination figure 303 can be defined as a predetermined shape, for example, an elliptical shape included inside the external circumscribing figure 302. In this case, the determination figure 303 is created by setting a ratio t (0 < t < 1) of the long axis 304 of the ellipse to the long side of the external circumscribing figure 302 and a ratio u (0 < u < 1) of the short axis 305 of the ellipse to the short side of the external circumscribing figure 302. In the determination figure database 130 of FIG. 6, initial values of the above-described ratios t and u are stored in correspondence with each character (alphabet, kana, Chinese character, etc.) in order to make the determination figure a predetermined elliptical shape.
[0061] As described in Patent Document 1, the judgment figures 303a and 303b are created so that the display characters 301a and 301b can be distinguished and seen when they do not overlap. Therefore, for multiple characters included in a character string displayed overlapping each other, the judgment figures 303 of both characters are similarly deformed from the initial value shapes according to the above initial values, and the similarity ratio when the two characters come into contact can be learned as a legible actual value.
[0062] 6, the shapes of the superimposed judgment figures 303a and 303b of "A" and "B" are determined so that the ellipses obtained by similarly deforming the respective ellipses according to the predetermined initial values of the ratios t and u (judgment figure database 130) are in contact with each other. The similarity ratio α1 of the judgment figure 303a and the similarity ratio α2 of the judgment figure 303b are learned as the actual values of the similarity ratio of the judgment figure 303 that make the superimposed display characters 301a and 301b legible.
[0063] In particular, by using the actual similarity ratio for the initial value shape set for each character, it is possible to learn the judgment figure for each character without having to learn it separately for each combination of multiple characters displayed on top of each other.
[0064] The model generation unit 120 executes the above-described judgment figure learning for each character whose initial value is stored in the judgment figure database 130, and determines a learning value αlrn of the similarity ratio for each character. The initial value of the learning value αlrn can be set to 1.0 for each character.
[0065] As a result, in a modified example of embodiment 1, the trained model generated by the model generation unit 120 can take each character displayed on top of each other as input, in addition to the input / output relationship (output 1 for inputs 1 to 4) described in embodiment 1, and output a trained value αlrn of the similarity ratio of the judgment figure for each character.
[0066] In this way, according to the learning device relating to the modified example of embodiment 1, in addition to the effects of embodiment 1, it is possible to generate a trained model that can determine with high accuracy whether multiple characters displayed on top of each other are legible by further using the judgment figure described in Patent Document 1.
[0067] In addition to the above-mentioned ellipse, the initial shape of the judgment figure can also be a circle or a rectangle. When the judgment figure is a rectangle or a circle, the learning value αlrn can be calculated from the ratio of the long side length of the judgment figure (rectangle) to the long side of the circumscribing figure (rectangle) and the ratio of the short side length of the judgment figure (rectangle) to the short side of the circumscribing figure. When the judgment figure is a circle, the learning value αlrn can be calculated from the ratio of the radius of the judgment figure to the radius of the circumscribing figure (circle).
[0068] Embodiment 2 In the second embodiment, an inference device and an image processing device to which the trained model according to the first embodiment is applied will be described.
[0069] FIG. 8 is a block diagram illustrating an example of the configuration of an inference device 200 and an image processing device 400 according to the second embodiment.
[0070] 8, the image processing device 400 includes an input unit 410, a display control unit 420, a display data storage unit 430, and a display unit 440. The input unit 410 specifies a display target for the display control unit 420. Specifically, a keyboard and a mouse of a PC (personal computer) operated by a user can be applied as the input unit 410. However, the input unit 410 is not limited to this example. The display control unit 420 and the display data storage unit 430 can also be realized by a program execution process by a CPU of the personal computer and a memory.
[0071] The display control unit 420 uses the inference device 200 to estimate whether a display object including multiple overlapping characters (character strings) specified by the input unit 410 is legible. In this case, the display control unit 420 extracts data indicating information on inputs 1 to 4 described in the first embodiment from the image data of the display object specified by the input unit 410. The extracted data indicating inputs 1 to 4 is provided to the inference device 200. Furthermore, the display control unit 420 has a display space for display on the display medium described above, and manages each character displayed in the display space using inputs 1, 2, 3, and 4. Here, the display space refers to a format consisting of a coordinate system, color data, and the like for display on a display medium (such as a display or printed matter) in a general system. The display control unit 420 then extracts inputs 1, 2, 3, and 4 of each character from the data in the display space and outputs them to the inference device 200. Furthermore, the display control unit 420 generates information related to the overlapping display of multiple characters included in the display object based on the output (inference result) of the inference device 200.
[0072] The display data storage unit 430 is a storage unit (memory) for storing data such as position and size required when displaying characters, figures, etc. on a medium, particularly for obtaining input 1, input 2, input 3, and input 4. From the viewpoint of simplifying the system, the display data storage unit 430 can store an index (pointer information) for accessing memory to obtain input 1, input 2, input 3, and input 4. Using the pointer information eliminates the need to copy the data of inputs 1 to 4 themselves, thereby reducing the memory capacity used. The display unit 440 is provided to display the information generated by the display control unit 420, and can be configured as a display or a printer for printing on paper.
[0073] The inference device 200 includes a data acquisition unit 210 and an inference unit 220. The functions of the data acquisition unit 210 and the inference unit 220 may also be realized by CPU processing of a computer shared with the image processing device 400, or by CPU processing of another computer that can communicate with the computer that constitutes the image processing device 400. As with the learning device 100 described in the first embodiment, the functions of the inference device 200 may also be realized on the cloud.
[0074] Similar to the data acquiring unit 110 in the first embodiment, the data acquiring unit 210 receives input 1 to input 4 for a plurality of characters included in a display target specified by the input unit 410, and generates model input data DIN, which will be input to the trained model described in the first embodiment. The model input data DIN is generated by the data acquiring unit 210 so as to have the same format as the training data xi in the first embodiment.
[0075] The inference unit 220 reads out the trained model created by the learning device according to embodiment 1 from the trained model storage unit 300, and generates the output 1 when the model input data DIN from the data acquisition unit 210 is input to the trained model as model output data DY.
[0076] A flowchart illustrating a first example of an inference process for generating model output data DY is shown in Figure 9. For example, the control process of Figure 9 can be executed by a computer on which a program for executing the functions of inference device 200 is installed.
[0077] 9, inference device 200 acquires, in S20, data relating to input 1 to input 4 for multiple characters to be displayed overlappingly and included in a display target specified by input unit 410. In S22, inference device 200 generates model input data DIN for each character from the data relating to input 1 to input 4 acquired in S20, and then, in S24, inputs the model input data DIN to the learned model stored in learned model storage unit 300.
[0078] In S26, inference device 200 obtains output 1 from the trained model for inputs 1 to 4. For example, output 1 indicates whether the combination of inputs 1 to 4 included in model input data DIN is included in the legible cluster or the unreadable cluster described above. In this case, model output data DY of inference unit 220 can be data indicating the determination result of whether each character displayed overlapping in the display object is legible or unreadable.
[0079] In this case, the display control unit 420 can control the display unit 440 to display the current display mode (corresponding to input 1 to input 4) of the display object specified by the input unit 410 and / or the drawing data (input 1 to input 4) and the judgment results (readable) of each character indicated by the model output data DY side by side to the user.
[0080] A flowchart illustrating a second example of inference processing for generating model output data DY is shown in Figure 10. The control processing in Figure 10 can also be executed by a computer on which a program for executing the functions of inference device 200 is installed.
[0081] Inference device 200 executes S20 to S26 similar to those in Fig. 9, and in S26 acquires a determination result of whether the text is legible or unreadable as output 1 of the trained model. In S27, inference device 200 branches the process depending on the determination result acquired in S26. If the determination result is "legible" (NO in S27), in S30, the determination result is output to display control unit 420 as model output data DY.
[0082] On the other hand, if the judgment result is "illegible" (YES judgment in S27), inference device 200 can perform a correction calculation on the image data in S28 to make it legible for at least one of inputs 1 to 4 included in the model input data DIN for each character, based on the correction amount described in embodiment 1. For example, in S28, the correction amounts for inputs 1 and 2 are calculated after inputs 3 and 4 of inputs 1 to 4 are set as constants.
[0083] Here, for inputs 1 to 4 acquired in S20, if (input 1, input 2, input 3, input 4) = (D1, D2, D3, D4) is written as input point P, then (input 1, input 2, input 3, input 4) = (D1P, D2P, D3, D4) can be defined as correction candidate point Q where inputs 3 and 4 are fixed at the above D3 and D4.
[0084] As described above, the trained model defines the boundary W between the legible and unreadable clusters, and the input point P is included in the unreadable cluster, while the correction candidate point Q is a point within the legible cluster. By fixing inputs 3 and 4, correction calculations can be performed using character attributes and colors as variables to ensure the visibility of overlaid characters under a given display medium and display format.
[0085] At this time, a correction vector (D1-D1P, D2-D2P, 0, 0) is defined as the difference between the input point P and the correction candidate point Q. That is, it can be understood that the correction vector is defined with the correction amounts calculated for each of inputs 1 to 4 as components.
[0086] For example, correction candidate point Q can be calculated as the point within the legible cluster where the magnitude of the correction vector is the smallest. In this case, correction candidate point Q can be the coordinate of the point on boundary W where input 3 and input 4 are D3 and D4. "Correction candidate data" can be obtained by correcting the current image data using corrected input 1 and input 2 included in correction candidate point Q.
[0087] When evaluating the magnitude of the correction vector, it is also possible to prioritize the items (variables) to be corrected by multiplying each of input 1 and input 2, or each of all data items included in input 1 and input 2 (in FIG. 3), by a weighting coefficient. Alternatively, it is also possible to treat some of all the data items described above as constants, like input 3 and input 4, and perform a correction calculation to find the correction candidate point Q. This allows for flexible correction calculations to be performed in accordance with the application to be displayed.
[0088] When the correction calculation in S28 is completed, inference device 200 outputs, in S29, the determination result indicating illegibility and the data of the correction candidate point calculated in S28 to display control unit 420 as model output data DY. In this case, display control unit 420 can control display unit 440 to display, side by side to the user, the current display mode (corresponding to input point P) and the display mode corresponding to correction candidate point Q for the display object specified by input unit 410, in addition to the determination result (illegible) of each character indicated by model output data DY.
[0089] Alternatively, when the determination in S27 is YES, that is, when it is determined that the display mode of the current inputs 1 to 4 is illegible, the display control unit 420 may notify the user using the display unit 440 that "visibility cannot be ensured" without executing the correction calculation in S28 and S29, and may output a message to the user urging them to input a correction instruction to the input unit 410. In this case, the current drawing data (inputs 1 to 4) may be output to the display unit 440, and a message or symbol informing the user that the data is "illegible" may be displayed near the overlapping portion inferred to be illegible, to prompt them to input a correction instruction.
[0090] In this way, the inference device and image processing device according to embodiment 2 can determine with high accuracy whether multiple overlapping characters included in a display target specified by a user are legible using the trained model according to embodiment 1. This makes it possible to arrange a large number of characters in a limited display area while ensuring human visibility by appropriately outputting display information for arranging overlapping character strings to the user.
[0091] Furthermore, if it is determined that the display mode specified in the current drawing data (input 1 to input 4) is unreadable, correction candidate points can be calculated using the trained model related to embodiment 1, and specific correction contents for the display mode can also be presented to the user.
[0092] A variation of the second embodiment. FIG. 11 is a block diagram illustrating an example of the configuration of an inference device 201 and an image processing device 401 according to a modification of the second embodiment.
[0093] 8, the inference device 201 includes a data acquisition unit 210 and an inference unit 221. The inference unit 221 includes a judgment figure creation unit 230 and a judgment unit 240. The functions of the judgment figure creation unit 230 and the judgment unit 240 can also be realized by CPU processing of a computer that is also used for the data acquisition unit 210 and the inference unit 220.
[0094] Image processing device 401 differs from image processing device 400 in Fig. 8 in that it includes a display control unit 421 instead of display control unit 420. Display control unit 421 performs the same operation as display control unit 420 using inference device 201. Other configurations and operations of image processing device 401 are the same as those of image processing device 400 according to embodiment 2, and therefore detailed description will not be repeated.
[0095] Inference device 201 infers whether the text is illegible or legible using the determination graphic described in the modified example of embodiment 1. The function of data acquisition unit 210 is similar to that described for inference device 200, and therefore detailed description will not be repeated.
[0096] The judgment figure creation unit 230 obtains a learned model such as a circle, ellipse, or rectangle from the learned model storage unit 300 based on inputs 1 to 4 obtained from the data acquisition unit 210, and fixes parameters for creating a judgment figure for each character. Here, "parameters" refer to attributes of the judgment figure; for example, if the judgment figure is a circle, the center coordinates and radius are fixed. Also, if the judgment figure is an ellipse, the center coordinates, long side length, and short side length are fixed, and if the judgment figure is a rectangle, the center coordinates, long side length, and short side length are fixed.
[0097] In this case, as explained in the modified example of the first embodiment, circumscribing figures 302 (FIGS. 6 and 7) are created for each of the multiple characters that are included in the display target specified by the input unit 410 and are displayed overlapping each other. Furthermore, a judgment figure 303 is created for each of the multiple characters by the above-mentioned parameter fixing using the learning value αlrn of the similarity ratio obtained by the learning process according to the modified example of the first embodiment. The judgment figure creation unit 230 creates a judgment figure 303 for each character and inputs it to the judgment unit 240.
[0098] The determination unit 240 determines whether the two characters displayed overlapping each other are "legible" or "illegible" based on the number of intersections between the determination figures 303 created by the determination figure creation unit 230. For example, the determination unit 240 can output the inference result RSLT such that it determines the character as "illegible" when the number of intersections between the determination figures 303 of the two characters is two or more, but determines the character as "legible" when the two determination figures 303 are in contact (the number of intersections is 1) or do not overlap (the number of intersections is 0).
[0099] Therefore, the display control unit 421 can highly accurately determine whether the display object is legible or not by using a determination figure determined (with fixed parameters) using a trained model.
[0100] In a modification of the second embodiment, when the determination in S27 of Fig. 10 is NO, that is, when the determination unit 240 determines that the image is "illegible," the display control unit 420 can notify the user that "visibility cannot be ensured" and output a message urging the user to input a correction instruction to the input unit 410, rather than performing the correction calculation (S28, S29).The flowchart of Fig. 10 can be modified so that the user is repeatedly prompted to input a correction instruction until there is no overlap between the determination figures 303 from inputs 1 to 4 after the correction instruction.
[0101] In this embodiment, when one character is displayed overlapping two or more characters, the above-described learning process and inference process are executed for each overlap between two of the characters. For example, when the character "A" and the character "B" are displayed overlapping, and further, the character "C" is displayed overlapping the character "B," the learning process using inputs 1 to 4 and the inference process for determining whether the overlapping characters are "readable" or "illegible" are executed for each of the overlapping characters "A" and "B" and the overlapping character "B" and "C."
[0102] In this embodiment, a case where unsupervised learning is applied to the learning algorithm used by the model generation unit and the inference unit has been described, but the present invention is not limited to this. That is, in addition to unsupervised learning, reinforcement learning, supervised learning, semi-supervised learning, etc. can also be applied to the learning algorithm. Furthermore, the learning algorithm used in the learning unit can be deep learning, which learns to extract feature amounts themselves, or any other known method.
[0103] Furthermore, when realizing unsupervised learning in this embodiment, it is possible to adopt other known clustering methods as a learning algorithm, not limited to the non-hierarchical clustering using the k-means method described above. For example, it is also possible to adopt hierarchical clustering such as the shortest distance method.
[0104] In addition, in this embodiment, the learning device and the inference device may be separate devices connected to the image processing device via a network, for example. Alternatively, the learning device and the inference device may be built into the image processing device. As described above, the learning device and the inference device may exist on a cloud server.
[0105] In the first embodiment, the model generation unit 120 may acquire training data from multiple image processing devices used for the same purpose, or may learn the output 1 using training data collected from multiple image processing devices used for different purposes. It is also possible to add or remove an image processing device that collects training data from the target during the process. Furthermore, a learning device that has learned the output 1 for a certain image processing device may be applied to another image processing device, and the output 1 may be re-learned and updated for the other image processing device.
[0106] The embodiments disclosed herein should be considered to be illustrative in all respects and not restrictive. The scope of the present disclosure is defined by the claims, not the above description, and is intended to include all modifications within the meaning and scope of the claims. [Explanation of symbols]
[0107] 10,20 display area, 100,101 learning device, 110,210 data acquisition unit, 120 model generation unit, 130 judgment figure database, 200,201 inference device, 220 inference unit, 230 judgment figure creation unit, 240 judgment unit, 300 learned model memory unit, 301a,301b display characters, 302,302a,302b circumscribed figures, 303,303a,303b judgment figures, 304 major axis, 305 minor axis, 400,401 image processing device, 410 input unit, 420,421 display control unit, 430 display data storage unit, 440 display unit, DIN model input data, DY model output data, Pa,Pb coordinates, xi learning data.
Claims
1. a data acquisition unit that acquires learning data for each of a plurality of characters displayed in an overlapping manner by combining a first input, a second input, a third input, and a fourth input from image data that indicates the display content of the character; a model generation unit that generates a trained model for inferring whether each of the characters is visible by distinguishing it from other characters displayed in an overlapping manner, from the first input, the second input, the third input, and the fourth input included in the training data; the first input includes information relating to a relative positional relationship between the plurality of characters and information relating to display attributes of the characters; the second input includes information related to a display color of the character; the third input includes information related to a display mode of the character; A learning device, wherein the fourth input includes information related to a display medium for the characters.
2. the information relating to the display attributes of the first input includes information relating to a font type and a size of the character; the second input includes information relating to a display color and a background color of the character and a display color of a character displayed superimposed on the character, the third input includes at least one of information as to whether the plurality of characters are to be displayed as a moving image or a still image, and information as to whether the plurality of characters are to be displayed as a two-dimensional image or a three-dimensional image, 2. The learning device according to claim 1, wherein the fourth input includes information indicating whether the display medium is a display or a printed material.
3. 3. The learning device according to claim 2, wherein the third input further includes information indicating that the character is to be displayed in a predetermined manner.
4. The learning device of claim 2, wherein the fourth input includes information indicating the resolution of the display when the display medium is the display, and information indicating the material and size of the printing surface when the display medium is the printed material.
5. The learned model uses the learning data including the first input, the second input, the third input, and the fourth input as input, and generates an output indicating a determination result of whether each of the characters can be visually distinguished from other characters displayed superimposed on the character when the plurality of characters are displayed according to the image data.
6. the trained model is configured to classify a combination of the first input, the second input, the third input, and the fourth input of each of the training data into a first cluster that is distinguishable from other characters displayed in an overlapping manner, and a second cluster that is indistinguishable from the other characters; 6. The learning device according to claim 5, wherein the output of the trained model further includes a correction amount of at least one of the first input, the second input, the third input, and the fourth input, for correcting the combination of the first input, the second input, the third input, and the fourth input so that the combination is classified into the first cluster when the combination is classified into the second cluster.
7. the model generation unit generates, for each of the characters, a circumscribing figure of a display area of the character in accordance with the image data and a determination figure for determining overlap between characters that are displayed in an overlapping manner; the determination figure is formed in a predetermined shape inside the circumscribed figure by referring to a database for each predetermined character, The learning device according to any one of claims 1 to 5, wherein the trained model is configured to learn the size of the judgment figure relative to the circumscribed figure when characters displayed on top of each other can be distinguished from each other.
8. a data acquisition unit that acquires model input data for each of a plurality of characters displayed in an overlapping manner by combining a first input, a second input, a third input, and a fourth input from image data that indicates display content of the character; an inference unit that outputs an inference result as to whether each of the plurality of characters displayed according to the image data is distinguishable and visible by inputting the model input data to a trained model that infers, from the first input, the second input, the third input, and the fourth input, whether each of the characters is distinguishable and visible from other characters displayed in an overlapping manner; the first input includes information relating to a relative positional relationship between the plurality of characters and information relating to display attributes of the characters; the second input includes information related to a display color of the character; the third input includes information related to a display mode of the character; The fourth input includes information related to the display medium of the character.
9. the trained model is configured to classify a combination of the first input, the second input, the third input, and the fourth input for each of the characters into a first cluster that is distinguishable from other characters displayed in an overlapping manner, and a second cluster that is indistinguishable from the other characters; The inference device of claim 8, wherein the inference unit outputs, as the inference result, a determination result as to whether each character is visible when the model input data is used as input to the trained model and when the plurality of characters are displayed according to the image data, the output of the trained model is distinguished from other characters displayed superimposed thereon.
10. the output of the trained model further includes a correction amount of at least any of the first input, the second input, the third input, and the fourth input, for correcting the combination of the first input, the second input, the third input, and the fourth input so that the combination is classified into the first cluster when the combination is classified into the second cluster; 10. The inference device according to claim 9, wherein the inference result output from the inference unit further includes correction candidate data obtained by correcting at least one of the first input, the second input, the third input, and the fourth input included in the model input data from the image data based on the correction amount.
11. The inference unit a judgment figure creation unit that creates a circumscribing figure of a display area of each of the characters according to the image data and a judgment figure that is formed in a predetermined shape inside the circumscribing figure and is used to determine overlap between characters that are displayed in an overlapping manner; a judgment unit that outputs the inference result based on the judgment figure created by the judgment figure creation unit, the learned model is configured to learn the size of the determination figure relative to the circumscribed figure when characters displayed in an overlapping manner are distinguishable from each other; the judgment diagram creation unit creates the judgment diagram for each of the plurality of characters by reflecting a learning result of the learned model; The inference device described in claim 8, wherein the judgment unit outputs the inference result indicating that each of the characters can be distinguished from other characters displayed overlapping the character when the judgment figure of each of the characters does not overlap with the judgment figures of the other characters.
12. the information relating to the display attributes of the first input includes information relating to a font type and a size of the character; the second input includes information relating to a display color and a background color of the character and a display color of a character displayed superimposed on the character, the third input includes at least one of information as to whether the plurality of characters are to be displayed as a moving image or a still image, and information as to whether the plurality of characters are to be displayed as a two-dimensional image or a three-dimensional image, 12. The inference device according to claim 8, wherein the fourth input includes information indicating whether the display medium is a display or a printed matter.
13. 13. The reasoning apparatus of claim 12, wherein the third input further includes information indicating that the character is to be displayed in a predetermined manner.
14. 13. The inference device of claim 12, wherein the fourth input includes information indicating the resolution of the display when the display medium is the display, and includes information indicating the material and size of the printing surface when the display medium is the printed material.
15. An inference device according to any one of claims 8 to 14; an input unit that designates the plurality of characters that are displayed in an overlapping manner as display targets; a display data storage unit for storing data including the first input, the second input, the third input, and the fourth input when the display object is displayed; an image processing device comprising: a display control unit that generates information on whether each of the plurality of characters is distinguishable and visible from the output of the inference device by inputting the first input, the second input, the third input, and the fourth input for each of the plurality of characters into the inference device.
Citation Information
Patent Citations
Neural network model construction method and system for recognizing image superimposed character area
CN110516665A
Method for processing character and picture
JP1990054297A
Document editing device
JP1992192069A
Character data processing system
JP1993080742A
Image processing device and image processing method
JP2011033705A