Font size detection method, text processing method and device
Patent Information
- Application Number
- CN202211567759.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-07
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-12-07
AI Technical Summary
[0027]本实施例提供的字号检测方法、文本处理方法及装置,包括:对待检测图像进行识别处理,得到待检测图像中的目标文本内容、以及目标文本内容占据的目标矩形区域,根据目标矩形区域确定目标文本内容的目标校准字号,根据目标文本内容和目标校准字号确定目标修正比例,并根据目标修正比例对目标校准字号进行修正得到目标字号,作为目标文本内容的字号检测结果,通过结合目标矩形区域确定目标校准字号,并基于目标修正比例对目标校准字号进行修正,得到目标文本内容的目标字号的技术特征,可以提高字号检测的准确性和可靠性。
Smart Images

Figure CN118155222B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing technology, and in particular to a font size detection method, a text processing method, and an apparatus. Background Technology
[0002] An application (APP) is software installed on a smartphone to improve upon the shortcomings of the original smartphone system and to personalize it.
[0003] The APP page includes text. Font size detection refers to detecting the font size of the text on the APP page. For example, the font size of the text on the APP page can be detected by means of Optical Character Recognition (OCR). Summary of the Invention
[0004] This disclosure provides a font size detection method, a text processing method, and an apparatus to improve the accuracy of font size detection.
[0005] In a first aspect, embodiments of this disclosure provide a font size detection method, including:
[0006] The image to be detected is processed to obtain the target text content in the image to be detected, and the target rectangular area occupied by the target text content;
[0007] Determine the target calibration font size of the target text content based on the target rectangular region;
[0008] A target correction ratio is determined based on the target text content and the target calibration font size, and the target calibration font size is corrected according to the target correction ratio to obtain the target font size, which is used as the font size detection result of the target text content.
[0009] Secondly, embodiments of this disclosure provide a text processing method, including:
[0010] The sample image is processed to obtain the sample text content in the sample image and the sample rectangular area occupied by the sample text content;
[0011] The sample calibration font size of the sample text content is determined based on the sample rectangular area;
[0012] A font size correction model is trained based on the sample text content and the sample calibrated font size, wherein the font size correction model is used to determine the target correction ratio of the font size of the target text content in the image to be detected.
[0013] Thirdly, embodiments of this disclosure provide a font size detection device, comprising:
[0014] The first recognition unit is used to perform recognition processing on the image to be detected to obtain the target text content in the image to be detected and the target rectangular area occupied by the target text content.
[0015] The first determining unit is used to determine the target calibration font size of the target text content based on the target rectangular region;
[0016] The second determining unit is used to determine the target correction ratio based on the target text content and the target calibration font size;
[0017] The correction unit is used to correct the target calibration font size according to the target correction ratio to obtain the target font size, which is used as the font size detection result of the target text content.
[0018] Fourthly, embodiments of this disclosure provide a text processing apparatus, including:
[0019] The second recognition unit is used to perform recognition processing on the sample image to obtain the sample text content in the sample image and the sample rectangular area occupied by the sample text content.
[0020] The third determining unit is used to determine the sample calibration font size of the sample text content based on the sample rectangular area;
[0021] The training unit is used to train a font size correction model based on the sample text content and the sample calibrated font size, wherein the font size correction model is used to determine the target correction ratio of the font size of the target text content in the image to be detected.
[0022] Fifthly, embodiments of this disclosure provide an electronic device, including: at least one processor and a memory;
[0023] The memory stores computer-executed instructions;
[0024] The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the font size detection method as described in the first aspect and various possible designs of the first aspect; or, causing the at least one processor to perform the text processing method as described in the second aspect and various possible designs of the second aspect.
[0025] Sixthly, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, wherein computer-executable instructions are stored therein, which, when executed by a processor, implement the font size detection method as described in the first aspect and various possible designs of the first aspect; or implement the text processing method as described in the second aspect and various possible designs of the second aspect.
[0026] In a seventh aspect, according to one or more embodiments of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the font size detection method as described in the first aspect and various possible designs of the first aspect; or, implements the text processing method as described in the second aspect and various possible designs of the second aspect.
[0027] The font size detection method, text processing method, and apparatus provided in this embodiment include: performing recognition processing on an image to be detected to obtain the target text content in the image to be detected and the target rectangular area occupied by the target text content; determining the target calibration font size of the target text content based on the target rectangular area; determining the target correction ratio based on the target text content and the target calibration font size; and correcting the target calibration font size based on the target correction ratio to obtain the target font size, which serves as the font size detection result of the target text content. By combining the target rectangular area to determine the target calibration font size and correcting the target calibration font size based on the target correction ratio, the technical characteristics of the target font size of the target text content are obtained, which can improve the accuracy and reliability of font size detection. Attached Figure Description
[0028] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0029] Figure 1 This is a schematic diagram illustrating the principle of a font size detection method according to an embodiment of the present disclosure;
[0030] Figure 2 This is a diagram illustrating the difference between the preset font size and the predicted font size.
[0031] Figure 3 A schematic diagram illustrating the differences in font height under the same font size settings in this disclosure;
[0032] Figure 4 This is a schematic diagram of a font size detection method according to an embodiment of the present disclosure;
[0033] Figure 5 This is a schematic diagram of a font size detection method according to another embodiment of the present disclosure;
[0034] Figure 6 This is a schematic diagram illustrating the principle of a font size detection method according to another embodiment of the present disclosure;
[0035] Figure 7 To adopt such Figure 6A schematic diagram of the verification results obtained by the method shown;
[0036] Figure 8 This is a schematic diagram illustrating the verification results of font size detection using the minimum bounding rectangle region.
[0037] Figure 9 This is a schematic diagram of a text processing method according to an embodiment of the present disclosure;
[0038] Figure 10 This is a schematic diagram of a text processing method according to another embodiment of the present disclosure;
[0039] Figure 11 This is a schematic diagram of a font size detection device according to an embodiment of the present disclosure;
[0040] Figure 12 This is a schematic diagram of a font size detection device according to another embodiment of the present disclosure;
[0041] Figure 13 This is a schematic diagram of a text processing apparatus according to an embodiment of the present disclosure;
[0042] Figure 14 This is a schematic diagram of a text processing apparatus according to another embodiment of the present disclosure;
[0043] Figure 15 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0044] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0045] Font size detection is a relatively basic capability in the content compliance detection of APP pages. It can be understood as detecting the font size of the text on the APP page in order to determine whether the font size of the text on the APP page meets the visual requirements based on the detected font size.
[0046] For example, considering the eyesight of middle-aged and elderly people, the font size (i.e., font size) of text in some APP pages is required to be relatively large. Accordingly, the preset font size in the page layout file can be set to a relatively large font size. However, there may be a difference between the actual font size of the text in the APP page and the preset font size in the page layout file. For example, the actual font size may be smaller than the preset font size. Therefore, before the APP is promoted and used, it is necessary to test the actual font size of the text in the APP page to meet the visual needs of APP users.
[0047] In some embodiments, the font size detection method may include the following steps:
[0048] Step 1: Take a screenshot of the app's page.
[0049] For example, the app can be run on a test device, and a screenshot of the app's screen can be taken, including text from the app's pages. The test device can be a smartphone or other device used for testing; this embodiment does not limit the scope.
[0050] The second step: Obtain the rectangular area (also called the area of the rectangle) to select the text to be detected from the screenshot of the screen.
[0051] For example, the content of an APP page includes text, which is the text to be detected. The text to be detected includes characters and numbers, etc. The rectangular area occupied by the text to be detected can be obtained through optical character recognition (OCR).
[0052] The third step: Determine the smallest bounding rectangle of the text to be detected within the rectangular region.
[0053] The rectangular area is used to select the area of the text to be detected. The rectangular area is actually slightly larger than the actual area of the text to be detected. In order to improve the accuracy of font size detection, the minimum bounding rectangle area can be obtained so that the area selected by the minimum bounding rectangle area is as close as possible to the actual area of the text to be detected.
[0054] like Figure 1 As shown, the text to be detected is “detect font size”. The area selected by the “rectangular area” of “detect font size” is larger than the actual area of “detect font size”. The area selected by the “minimum bounding rectangle area” is smaller than the area selected by the “rectangular area”. The area selected by the “minimum bounding rectangle area” is closer to the actual area of “detect font size”.
[0055] For example, the text may be in horizontal layout or vertical layout. Horizontal layout may also be referred to as horizontally arranged layout, and vertical layout may also be referred to as vertically arranged layout. Correspondingly, the layout of the text to be detected may be horizontal layout or vertical layout.
[0056] It should be understood that, Figure 1 the description is made by taking the horizontal layout as an example for illustrative demonstration, and cannot be construed as a limitation on the layout of the text to be detected, the minimum bounding rectangle region, and the like.
[0057] Fourth step: determining the predicted font size of the text to be detected according to the minimum bounding rectangle region.
[0058] For example, if the text to be detected is a text in horizontal layout, the height difference of the minimum bounding rectangle region may be determined, and the height difference is determined as the predicted font size of the text to be detected. If the text to be detected is a text in vertical layout, the width difference of the minimum bounding rectangle region may be determined, and the width difference is determined as the predicted font size of the text to be detected.
[0059] However, the preset font size may not correspond to the minimum bounding rectangle region, for example Figure 2 as shown, the preset font size is 60 points (pt), while the predicted font size obtained based on the method of the foregoing steps 1 to 4 is 53 pt.
[0060] Moreover, different texts to be detected exhibit different font height differences under the same font size setting. For example Figure 3 as shown, "我", "w", and "0" have the same preset font size in the page layout file, but the height difference of the minimum bounding rectangle region of "我" is the largest, followed by that of "0", and the height difference of the minimum bounding rectangle region of "w" is the smallest.
[0061] In order to improve the accuracy of font size detection, the present disclosure provides a technical concept obtained through creative work: determining a target calibrated font size of target text content according to a target rectangular region occupied by the target text content in an image to be detected, determining a target correction proportion based on the target text content and the target calibrated font size, correcting the target calibrated font size based on the target correction proportion to obtain a target font size, and using the target font size as a font size detection result of the target text content.
[0062] The technical solutions of the present disclosure and how the technical solutions of the present disclosure solve the above technical problems are described in detail below with specific embodiments. The several specific embodiments described below may be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. Embodiments of the present disclosure will be described below with reference to the accompanying drawings.
[0063] Please refer to Figure 4 , Figure 4 This is a schematic diagram of a font size detection method according to an embodiment of the present disclosure.
[0064] like Figure 4 As shown, the method includes:
[0065] S401: Perform recognition processing on the image to be detected to obtain the target text content in the image to be detected and the target rectangular area occupied by the target text content.
[0066] The execution subject in this embodiment can be a font size detection device, which can be a server (such as a local server, a cloud server, or a server cluster), a computer, a terminal device, a processor, a chip, etc. This embodiment does not limit the device.
[0067] The image to be detected is an image containing text content. To facilitate the differentiation of the text content in the image to be detected from other text content, such as the text content in the sample images later, the text content in the image to be detected is referred to as the target text content. Similarly, the rectangular area occupied by the target text content is referred to as the target rectangular area.
[0068] S402: Determine the target calibration font size of the target text content based on the target rectangular area.
[0069] The target calibration font size is the font size of the target text content determined based on the target rectangular area.
[0070] S403: Determine the target correction ratio based on the target text content and the target calibration font size, and correct the target calibration font size according to the target correction ratio to obtain the target font size, which is used as the font size detection result of the target text content.
[0071] The target correction ratio is the ratio used to correct the target calibration font size. The target font size is the font size after correcting the font size of the target text content.
[0072] For example, in this embodiment, the approximate font size of the target text content (i.e., the target calibration font size) can be determined first based on the target rectangular area, and then the approximate font size (i.e., the target calibration font size) can be corrected based on the target correction ratio to obtain a highly accurate font size (i.e., the target font size).
[0073] Based on the above analysis, this disclosure provides a font size detection method, including: performing recognition processing on an image to be detected to obtain the target text content in the image to be detected and the target rectangular area occupied by the target text content; determining the target calibration font size of the target text content based on the target rectangular area; determining the target correction ratio based on the target text content and the target calibration font size; and correcting the target calibration font size based on the target correction ratio to obtain the target font size, which serves as the font size detection result of the target text content. In this embodiment, by combining the target rectangular area to determine the target calibration font size and correcting the target calibration font size based on the target correction ratio, the technical characteristics of the target font size of the target text content are obtained, which can improve the accuracy and reliability of font size detection.
[0074] To enable readers to gain a deeper understanding of the implementation principles of this disclosure, the following is combined with... Figure 5 The implementation principles of this disclosure will be explained in more detail. Among them, Figure 5 This is a schematic diagram of a font size detection method according to another embodiment of the present disclosure.
[0075] like Figure 5 As shown, the method includes:
[0076] S501: Perform recognition processing on the image to be detected to obtain the target text content in the image to be detected and the target rectangular area occupied by the target text content.
[0077] It should be understood that, in order to avoid tedious descriptions, the same technical features as those in the above embodiments will not be repeated in this embodiment.
[0078] For example, the implementation principle of S501 can be found in the implementation principle of S401, which will not be repeated here.
[0079] S502: Determine the target calibration font size of the target text content based on the target rectangular area.
[0080] For example, the implementation principle of S502 can be found in the implementation principle of S402, which will not be repeated here.
[0081] In some embodiments, the target calibration font size is determined based on the target boundary information of the target rectangular region.
[0082] Among them, the target boundary information can be understood as the pixel coordinates of the boundary of the target text content in the target rectangular area, that is, the two-dimensional coordinates based on the image coordinate system.
[0083] In this embodiment, by determining the target calibration font size based on the target boundary information, the target calibration font size can be closely matched with the boundary of the target text content, so that the target calibration font size can relatively accurately represent the font size of the target text content, thereby improving the accuracy and reliability of the target calibration font size.
[0084] In some embodiments, if the text in the image to be detected is in a horizontal layout, the target boundary information includes the upper and lower boundary information of the text in the image to be detected; if the text in the image to be detected is in a vertical layout, the target boundary information includes the left and right boundary information of the text in the image to be detected.
[0085] In this embodiment, different target boundary information is used to determine the target calibration font size for text in different layouts of images to be detected, which can achieve flexibility and diversity in determining the target calibration font size.
[0086] In some embodiments, if the text in the image to be detected is in a horizontal format, the target boundary information is determined by projecting and superimposing the image in the vertical direction of the text in the image to be detected, obtaining the target gray value, and then determining the target boundary information based on the target gray value.
[0087] If the text in the image to be detected is in a vertical format, the target boundary information is determined by projecting and overlaying the image to be detected horizontally based on the horizontal direction of the text in the image to obtain the target gray value, and then determining the target boundary information based on the target gray value.
[0088] In other words, target boundary information can be determined by projection overlay processing. The direction of projection overlay processing is different for different layouts of text in the image to be detected. For text in a horizontal layout image to be detected, the image to be detected is projected and overlaid based on the vertical direction. For text in a vertical layout image to be detected, the image to be detected is projected and overlaid based on the horizontal direction.
[0089] For example, taking the text in the image to be detected as horizontally formatted text as an example, determining the target boundary information may include the following steps:
[0090] The first step is to project and overlay the image to be detected along the vertical direction (i.e., the vertical direction of the image to be detected, which is the y-axis direction of the image to be detected) to obtain the target gray value L.
[0091] The second step involves traversing the target grayscale values from top to bottom until the top threshold (i.e., upper boundary information) of the text in the image to be detected is determined based on the current grayscale value and the previous grayscale value. The top threshold represents the coordinates of the top of the text in the image to be detected, relative to the image coordinate system of the image to be detected.
[0092] For example, i is a counter that starts from 0 and increments sequentially to calculate k until k is greater than a preset threshold. Here, k = abs(L(i+1)-L(i)) / L(i), where abs represents the absolute value.
[0093] In contrast, in the image to be detected, the gray value of the target without text will be higher, while the gray value of the target with text will be lower. Therefore, the top threshold value can be determined based on the difference between the target gray values of the two consecutive traversals.
[0094] The third step is to traverse the target grayscale values from bottom to top until the bottom boundary information (i.e., lower boundary information) of the text in the image to be detected is determined based on the current grayscale value and the previous grayscale value. The bottom boundary information represents the coordinates of the bottom of the text in the image to be detected relative to the image coordinate system of the image to be detected.
[0095] Accordingly, the target calibration font size = bottom critical value - top critical value.
[0096] The implementation principle for determining the bottom critical value can be found in the implementation principle for determining the top critical value, and will not be repeated here. Furthermore, the implementation principle for determining the target boundary information of the vertical layout is the same as that for determining the target boundary information of the horizontal layout, and will not be repeated here.
[0097] In this embodiment, the target grayscale value is obtained through projection overlay processing, and the target boundary information is determined based on the target grayscale value. The difference between the target grayscale values with and without text in the image to be detected is fully considered, so that the determined target boundary information can relatively accurately represent the boundary of the text in the image to be detected. Therefore, when the target calibration font size is determined based on the target boundary information, the target calibration font size can have high accuracy and reliability.
[0098] S503: Encode the target text content and the target calibration font size separately to obtain the target text encoding vector of the target text content and the target font size encoding vector of the target calibration font size.
[0099] In this embodiment, the encoding process is not limited.
[0100] S504: Determine the target correction ratio based on the target text encoding vector and the target font size encoding vector.
[0101] In some embodiments, S504 includes: performing a fusion process on the target text encoding vector and the target font size encoding vector to obtain a target fusion vector, and determining a target correction ratio based on the target fusion vector.
[0102] For example, the target font size encoding vector and the target text encoding vector can be concatenated to obtain the target fusion vector. In the target fusion vector, the positions of the target font size encoding vector and the target text encoding vector are not limited. The target font size encoding vector can be placed first, or the target text encoding vector can be placed first.
[0103] Alternatively, a target text encoding vector can be inserted into the target font size encoding vector to obtain the target fusion vector. Another method is to insert a target font size encoding vector into the target text encoding vector to obtain the target fusion vector, and so on. These methods will not be listed here.
[0104] In some embodiments, before fusing the target font size encoding vector and the target text encoding vector, the target text encoding vector can be regularized, such as by performing L1 or L2 regularization, to eliminate the influence of the text length of the target text content and improve the effectiveness and stability of the fused vector.
[0105] In this embodiment, by fusion (concat) the target font size encoding vector and the target text encoding vector, the resulting target fused vector can have features in both the font size dimension and the text content dimension, thereby improving the effectiveness and reliability of font size detection.
[0106] In some embodiments, determining the target correction ratio based on the target fusion vector may include: inputting the target fusion vector into a pre-trained font size correction model and outputting the target correction ratio.
[0107] Among them, the font size correction model is a model trained on sample images to predict the correction ratio of the font size.
[0108] The training principle of the font size correction model can be found in the description of the following embodiments, and will not be repeated here.
[0109] In other words, in some embodiments, the target correction ratio can be determined by combining a network model, thereby improving the diversity and flexibility in determining the target correction ratio.
[0110] In some embodiments, encoding the target calibration font number to obtain the target font number encoding vector of the target calibration font number includes the following steps:
[0111] First step: Determine the target grouping interval corresponding to the target calibration font size based on the preset mapping relationship between font size and font size grouping interval.
[0112] For example, the mapping relationship is used to characterize the correspondence between font size and font size grouping intervals, which can be determined based on requirements, historical records, and experiments.
[0113] The second step is to encode the target group interval to obtain the target font size encoding vector.
[0114] For example, the font size grouping ranges include five groups: (20pt-40pt), (40pt-60pt), (60pt-80pt), (80pt-100pt), and (100pt-120pt).
[0115] If the target calibration font size is 46pt, then the target calibration font size belongs to the second group (40pt-60pt), that is, (40pt-60pt) is the target grouping interval. Encoding the target grouping interval can be understood as encoding the index of the target grouping interval (i.e., the second group), to obtain the target font size encoding vector [0, 1, 0, 0, 0].
[0116] S505: Correct the target calibration font size according to the target correction ratio to obtain the target font size, which is used as the font size detection result of the target text content.
[0117] For example, regarding the implementation principle of S505, please refer to part of the implementation principle in S403, which will not be repeated here.
[0118] To help readers gain a deeper understanding of the implementation principles of this embodiment, the following is a combination of... Figure 6 This embodiment will now be described in more detail.
[0119] like Figure 6 As shown, optical character recognition is performed on the image to be detected, resulting in the following: Figure 6 The target rectangular region shown includes, for example, the target rectangular region. Figure 6 The target text content shown is "Text".
[0120] By employing the method described in the above embodiments, the following can be obtained: Figure 6 The target calibration font size y of the target text content "Text" shown is as follows: Figure 6 The "target font size encoding vector" corresponding to the target calibration font size shown in the figure is as follows: Figure 6 The "target text encoding vector" of the target text content "Text" shown is, for example... Figure 6 The "L1 norm" of the "target text encoding vector" shown will not be elaborated here.
[0121] Similarly, the "target font size encoding vector" and the "L1 norm" can be fused to obtain, as shown below. Figure 6 The “target fusion vector” is shown.
[0122] like Figure 6As shown, the “target fusion vector” is input into the “font size correction model”, and the “target correction ratio” of the target text content “Text” is output.
[0123] like Figure 6 As shown, the product of "target correction ratio" and "target calibration font size y" is determined as the target font size of the target text content "Text".
[0124] In some embodiments, in order to make the APP meet the user's visual needs, the target font size in the APP's page layout file can be adjusted based on the target font size.
[0125] Based on the above analysis, the font size grouping intervals can include five groups: (20pt-40pt), (40pt-60pt), (60pt-80pt), (80pt-100pt), and (100pt-120pt). To verify the reliability of the font size detection method used in this embodiment, the font size detection results for each group of intervals were verified. The verification results can be found in [reference needed]. Figure 7 .
[0126] in, Figure 7 The horizontal axis represents the error, that is, the error between the target font size predicted by the font size detection method of this embodiment and the actual font size, and the vertical axis represents the proportion.
[0127] Figure 8 The results of font size detection using the minimum bounding rectangle area are shown. The horizontal axis represents the error, which is the error between the target font size calculated using the minimum bounding rectangle area and the actual font size. The vertical axis represents the proportion.
[0128] Combination Figure 7 and Figure 8 It can be seen that the target font size predicted by the font size detection method of this embodiment is closer to the actual font size. That is, the font size detection method of this embodiment can be more accurate, thus obtaining a more accurate target font size.
[0129] Please see Figure 9 , Figure 9 This is a schematic diagram of a text processing method according to an embodiment of the present disclosure.
[0130] like Figure 9 As shown, the method includes:
[0131] S901: Perform recognition processing on the sample image to obtain the sample text content in the sample image and the sample rectangular area occupied by the sample text content.
[0132] The execution subject in this embodiment can be a text processing device. The text processing device can be the same as the font size detection device or a different device. This embodiment does not limit the type of device.
[0133] If the text processing device is a different device from the font size detection device, then when the text processing device trains the font size correction model, it can transmit the font size correction model to the font size detection device.
[0134] For example, the sample image can be a screenshot of a screen page as described in the above embodiments. To improve the effectiveness and reliability of text processing, the sample image can be obtained by taking a screenshot of the actual device.
[0135] For example, running an app on a feature phone and capturing an image of the page while the app is running is called a sample image.
[0136] This embodiment does not limit the number of sample images or the text content of the sample images; these can be determined based on requirements, historical records, and experiments.
[0137] Accordingly, the sample text content is the content of the APP page. The sample text content can be one or more of the following: Chinese characters, numbers, English letters, and punctuation marks. The font of the sample text content can be the system default font of the feature phone, or other fonts, and the font size can be any size between 20pt and 120pt.
[0138] In some embodiments, after obtaining a corresponding number (e.g., 100,000) of sample images, the sample images can be enriched, such as by using a certain proportion of real background textures and / or randomly switching background colors to obtain richer sample data.
[0139] For example, 1% of sample images are randomly selected from 100,000 sample images, and the background of the selected sample images is replaced with a real background texture (the real background textures replaced by different sample images can be the same or different), thus obtaining the newly added 1% of sample images.
[0140] Similarly, 1% of the sample images are randomly selected from 100,000 sample images, and the background color of the selected sample images is replaced with a certain background color (the background color replaced by different sample images can be the same or different, and the background color can be a solid color or a gradient color, etc.), thus obtaining the newly added 1% of sample images.
[0141] The sample rectangular area can be understood as using optical character recognition to process the sample image, thereby obtaining a rectangular area for selecting the sample text content, which includes the sample text content.
[0142] S902: Determine the sample calibration font size of the sample text content based on the sample rectangular area.
[0143] S903: A font size correction model is trained based on sample text content and sample calibrated font size. This model is used to determine the target correction ratio for the font size of the target text content in the image to be detected.
[0144] In this embodiment, the font size correction model is trained by combining the sample calibration font size and the sample text content. By combining multiple dimensions of content to train the font size correction model, the font size correction model can have high accuracy and reliability, and fully consider the actual font size of the sample text content, so that the font size correction model can better fit the actual scenario and meet the prediction needs of the actual scenario.
[0145] To enable readers to gain a deeper understanding of the implementation principles of this disclosure, the following is combined with... Figure 10 The implementation principles of this disclosure will be explained in more detail. Among them, Figure 10 This is a schematic diagram of a text processing method according to another embodiment of the present disclosure.
[0146] like Figure 10 As shown, the method includes:
[0147] S1001: Perform recognition processing on the sample image to obtain the sample text content in the sample image and the sample rectangular area occupied by the sample text content.
[0148] Similarly, the technical features that are the same as those in the above embodiments will not be repeated in this embodiment.
[0149] For example, the implementation principle of S1001 can be found in the description of S901, which will not be repeated here.
[0150] S1002: Obtain the sample boundary information of the sample text content based on the sample rectangular area, and determine the sample calibration font size of the sample text content based on the sample boundary information.
[0151] Among them, the sample boundary information can be understood as the pixel coordinates of the boundary of the sample text content in the sample rectangular area, that is, the two-dimensional coordinates based on the image coordinate system.
[0152] Similarly, in this embodiment, by first determining the sample boundary information and then combining it with the sample boundary information to determine the sample calibration font size, the sample calibration font size can be made to closely match the boundary of the sample text content. This allows the sample calibration font size to represent the actual font size of the sample text content relatively accurately, thereby improving the accuracy and reliability of the sample calibration font size.
[0153] In some embodiments, if the text in the sample image is in a horizontal layout, the sample boundary information includes the upper and lower boundary information of the text in the sample image; if the text in the sample image is in a vertical layout, the sample boundary information includes the left and right boundary information of the text in the sample image.
[0154] Similarly, in this embodiment, different sample boundary information is used to determine the sample calibration font size for text in sample images of different layouts, which can achieve flexibility and diversity in determining the sample calibration font size.
[0155] In some embodiments, if the text in the sample image is in a horizontal format, the sample boundary information is determined by projecting and overlaying the sample image vertically based on the vertical direction of the text in the sample image to obtain the sample grayscale value, and then determining the grayscale value based on the sample grayscale value.
[0156] If the text in the sample image is in a vertical format, the sample boundary information is determined by projecting and overlaying the sample image horizontally based on the horizontal direction of the text in the sample image to obtain the sample grayscale value, and then determining the grayscale value based on the sample grayscale value.
[0157] In other words, the sample boundary information can be determined by projection overlay processing. The direction of projection overlay processing is different for the text in sample images of different layouts. For the text in the horizontal layout sample image, the sample image is projected and overlaid based on the vertical direction, while for the text in the vertical layout sample image, the sample image is projected and overlaid based on the horizontal direction.
[0158] In this embodiment, sample grayscale values are obtained through projection overlay processing. Sample boundary information is determined based on the sample grayscale values. The difference between sample grayscale values with and without sample text content is fully considered so that the determined sample boundary information can relatively accurately represent the boundary of the sample text content. Therefore, when determining the sample calibration font size based on the sample boundary information, the sample calibration font size can have high accuracy and reliability.
[0159] S1003: Determine the sample grouping interval corresponding to the sample calibration font size based on the preset mapping relationship between font size and font size grouping interval.
[0160] S1004: Encode the sample grouping intervals to obtain the sample font size encoding vector.
[0161] For example, the mapping relationship is used to characterize the correspondence between font size and font size grouping intervals, which can be determined based on requirements, historical records, and experiments. For example, the font size grouping intervals include five groups: (20pt-40pt), (40pt-60pt), (60pt-80pt), (80pt-100pt), and (100pt-120pt).
[0162] If the calibration font size is 46pt, then the sample calibration font size belongs to the second group (40pt-60pt), that is, (40pt-60pt) is the sample grouping interval. Encoding the sample grouping interval can be understood as encoding the index of the sample grouping interval (i.e., the second group), resulting in the sample font size encoding vector [0, 1, 0, 0, 0].
[0163] S1005: Encode the sample text content to obtain the sample text encoding vector, and fuse the font size encoding vector and the text encoding vector to obtain the fused vector.
[0164] This embodiment does not limit the encoding process. For example, it can encode multi-class labels to obtain multi-class label (one-hot) vectors (i.e., sample text encoding vectors).
[0165] For example, the sample font size encoding vector and the sample text encoding vector can be concatenated to obtain the sample fusion vector. In the sample fusion vector, the positions of the sample font size encoding vector and the sample text encoding vector are not limited. The sample font size encoding vector can be placed first, or the sample text encoding vector can be placed first.
[0166] Alternatively, sample text encoding vectors can be inserted into sample font size encoding vectors to obtain sample fusion vectors. Sample font size encoding vectors can also be inserted into sample text encoding vectors to obtain sample fusion vectors, and so on. These will not be listed here.
[0167] In some embodiments, before fusing the sample font size encoding vector and the sample text encoding vector, the sample text encoding vector can be regularized, such as by performing L1 or L2 regularization, to eliminate the influence of the text length of the sample text content and improve the effectiveness and stability of the fused sample vector.
[0168] S1006: Generate a font size correction model based on the sample fusion vector and the sample calibration font size. The font size correction model is used to determine the target correction ratio for the font size of the target text content in the image to be detected.
[0169] In this embodiment, by fusion (concat) the sample font size encoding vector and the sample text encoding vector, the resulting sample fusion vector can have features in both the font size dimension (specifically, the actual font size) and the text content dimension. Thus, when the font size correction model is generated by combining the sample fusion vector, the font size correction model can learn both the features in the font size dimension and the features in the text content dimension, thereby improving the accuracy and reliability of the font size correction model.
[0170] In some embodiments, S1006 may include the following steps:
[0171] The first step is to predict the prediction ratio based on the sample fusion vector. The prediction ratio represents the font size ratio between the preset standard font size of the predicted sample text content and the sample calibrated font size.
[0172] The second step is to calculate the true ratio between the preset standard font size and the sample calibration font size.
[0173] The third step: Generate a font size correction model based on the predicted ratio and the actual ratio.
[0174] For example, the basic network model can adopt a single-layer, fully connected structure (such as including linear layers), and the activation function can be the exp activation function.
[0175] Based on the above analysis, it can be seen that the sample font size encoding vector and the sample text encoding vector can be fused to obtain the sample fusion vector, which can then be input into the basic network model.
[0176] In other embodiments, the base network model can be a multi-layer neural network structure (such as a structure with multiple fully connected layers). There is no need to fuse the sample font size encoding vector and the sample text encoding vector. The sample font size encoding vector and the sample text encoding vector can be connected to their respective fully connected layers. Finally, the outputs of the tanh activation function (such as the activation function after the sample font size encoding vector passes through its corresponding fully connected layer) and the exp activation function (such as the activation function after the sample text encoding vector passes through its corresponding fully connected layer) are added together to generate the font size correction model (in order to decouple the correlation between the sample font size encoding vector and the sample text encoding vector).
[0177] The third step may include: determining the loss function between the predicted and actual proportions, and iteratively training the font size correction model based on the loss function. The loss function can be implemented using mean absolute error (MAE).
[0178] In this embodiment, the true ratio is obtained by calculation, and the predicted ratio is obtained by prediction, in order to combine the true ratio with the predicted ratio.
[0179] The font size correction model is trained using both the actual and predicted proportions. This allows the model to gradually satisfy the condition that the difference between the predicted and actual proportions decreases, thereby improving the accuracy of the font size correction model.
[0180] Sex and reliability.
[0181] According to embodiments of this disclosure, this disclosure also provides a font size detection device.
[0182] Please see Figure 11 , Figure 11 This is a schematic diagram of a font size detection device according to an embodiment of the present disclosure, as shown below.
[0183] Figure 11 As shown, the font size detection device 1100 includes:
[0184] The first recognition unit 1101 is used to perform recognition processing on the image to be detected to obtain the image to be detected.
[0185] The target text content and the target rectangular area occupied by the target text content.
[0186] The first determining unit 1102 is used to determine the target calibration font size of the target text content based on the target rectangular area.
[0187] The second determining unit 1103 is used to determine the target correction ratio based on the target text content and the target calibration font size.
[0188] The correction unit 1104 is used to correct the target calibration font size according to the target correction ratio to obtain the target font size, which is used as the font size detection result of the target text content.
[0189] Please see Figure 12 , Figure 12 This is a schematic diagram of a font size detection device according to another embodiment of the present disclosure, as shown below.
[0190] Figure 12 As shown, the font size detection device 1200 includes:
[0191] The first recognition unit 1201 is used to perform recognition processing on the image to be detected to obtain the image to be detected.
[0192] The target text content and the target rectangular area occupied by the target text content.
[0193] The first determining unit 1202 is used to determine the target calibration font size of the target text content based on the target rectangular area.
[0194] The second determining unit 1203 is used to determine the target correction ratio based on the target text content and the target calibration font size.
[0195] In some embodiments, combined with Figure 12 It can be seen that the second determining unit 1203 includes:
[0196] The first encoding subunit 12031 is used to encode the target text content and the target calibration font size respectively, so as to obtain the target text encoding vector of the target text content and the target font size encoding vector of the target calibration font size.
[0197] 0 In some embodiments, the first coding subunit 12031 includes:
[0198] The second determining module is used to determine the target grouping interval corresponding to the target calibration font size based on the preset mapping relationship between font size and font size grouping interval.
[0199] The encoding module is used to encode the target group interval to obtain the target font size encoding vector.
[0200] Subunit 12032 is defined to determine the target correction ratio based on the target text encoding vector and the target font size encoding vector.
[0201] In some embodiments, determining subunit 12032 includes:
[0202] The fusion module is used to fuse the target text encoding vector and the target font size encoding vector to obtain the target fused vector.
[0203] The first determining module is used to determine the target correction ratio based on the target fusion vector.
[0204] In some embodiments, the first determining module is used to input the target fusion vector into a pre-trained font size correction model and output the target correction ratio.
[0205] Among them, the font size correction model is a model trained on sample images to predict the correction ratio of the font size.
[0206] The correction unit 1204 is used to correct the target calibration font size according to the target correction ratio to obtain the target font size, which is used as the font size detection result of the target text content.
[0207] In some embodiments, the target calibration font size is determined based on the target boundary information of the target rectangular region.
[0208] In some embodiments, if the text in the image to be detected is in a horizontal layout, the target boundary information includes the upper and lower boundary information of the text in the image to be detected; if the text in the image to be detected is in a vertical layout, the target boundary information includes the left and right boundary information of the text in the image to be detected.
[0209] In some embodiments, if the text in the image to be detected is in a horizontal format, the target boundary information is determined by projecting and superimposing the image in the vertical direction of the text in the image to be detected, obtaining the target gray value, and then determining the target boundary information based on the target gray value.
[0210] If the text in the image to be detected is in a vertical format, the target boundary information is determined by projecting and overlaying the image to be detected horizontally based on the horizontal direction of the text in the image to obtain the target gray value, and then determining the target boundary information based on the target gray value.
[0211] According to embodiments of this disclosure, this disclosure also provides a text processing apparatus.
[0212] Please see Figure 13 , Figure 13 This is a schematic diagram of a text processing apparatus according to an embodiment of the present disclosure, as shown below. Figure 13 As shown, the text processing device 1300 includes:
[0213] The second recognition unit 1301 is used to perform recognition processing on the sample image to obtain the sample text content in the sample image and the sample rectangular area occupied by the sample text content.
[0214] The third determining unit 1302 is used to determine the sample calibration font size of the sample text content based on the sample rectangular area.
[0215] Training unit 1303 is used to train a font size correction model based on sample text content and sample calibration font size, wherein the font size correction model is used to determine the target correction ratio of the font size of the target text content in the image to be detected.
[0216] Please see Figure 14 , Figure 14 This is a schematic diagram of a text processing apparatus according to another embodiment of the present disclosure, such as... Figure 14 As shown, the text processing device 1400 includes:
[0217] The second recognition unit 1401 is used to perform recognition processing on the sample image to obtain the sample text content in the sample image and the sample rectangular area occupied by the sample text content.
[0218] The third determining unit 1402 is used to determine the sample calibration font size of the sample text content based on the sample rectangular area.
[0219] Training unit 1403 is used to train a font size correction model based on sample text content and sample calibrated font size. The font size correction model is used to determine the target correction ratio of the font size of the target text content in the image to be detected.
[0220] Combination Figure 14 It is understood that, in some embodiments, the training unit 1403 includes:
[0221] The second encoding subunit 14031 is used to encode the sample text content and the sample calibration font size respectively, to obtain the sample text encoding vector of the sample text content and the sample font size encoding vector of the sample calibration font size.
[0222] The fusion subunit 14032 is used to fuse the sample font size encoding vector and the sample text encoding vector to obtain the sample fusion vector.
[0223] Generating subunit 14033 is used to generate a font size correction model based on the sample fusion vector and the sample calibration font size.
[0224] In some embodiments, generating subunit 14033 includes:
[0225] The prediction module is used to predict the prediction ratio based on the sample fusion vector. The prediction ratio is used to characterize the ratio between the preset standard font size of the predicted sample text content and the font size of the sample calibration font size.
[0226] The calculation module is used to calculate the true ratio between the preset standard font size and the sample calibration font size.
[0227] The generation module is used to generate a font size correction model based on the predicted ratio and the actual ratio.
[0228] In some embodiments, the generation module includes:
[0229] The determination submodule is used to determine the sample grouping interval corresponding to the sample calibration font size based on the preset mapping relationship between font size and font size grouping interval.
[0230] The encoding submodule is used to encode the sample grouping intervals to obtain the sample font size encoding vector.
[0231] In some embodiments, the standard font size of the sample is determined based on the sample boundary information of the sample rectangular region.
[0232] In some embodiments, if the text in the sample image is in a horizontal layout, the sample boundary information includes the upper and lower boundary information of the text in the sample image; if the text in the sample image is in a vertical layout, the sample boundary information includes the left and right boundary information of the text in the sample image.
[0233] In some embodiments, if the text in the sample image is in a horizontal format, the sample boundary information is determined by projecting and overlaying the sample image vertically based on the vertical direction of the text in the sample image to obtain the sample grayscale value, and then determining the grayscale value based on the sample grayscale value.
[0234] If the text in the sample image is in a vertical format, the sample boundary information is determined by projecting and overlaying the sample image horizontally based on the horizontal direction of the text in the sample image to obtain the sample grayscale value, and then determining the grayscale value based on the sample grayscale value.
[0235] According to embodiments of this disclosure, this disclosure also provides an electronic device and a readable storage medium.
[0236] According to embodiments of this disclosure, this disclosure also provides a computer program product, the program product comprising: a computer program stored in a readable storage medium, at least one processor of an electronic device being able to read the computer program from the readable storage medium, and the at least one processor executing the computer program causing the electronic device to perform the scheme provided in any of the above embodiments.
[0237] refer to Figure 15 The diagram illustrates a structural schematic of an electronic device 1500 suitable for implementing embodiments of the present disclosure. The electronic device 1500 can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, personal digital assistants (PDAs), portable Android devices (PADs), portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 15 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0238] like Figure 15As shown, the electronic device 1500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 1501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1502 or a program loaded from a storage device 1508 into a random access memory (RAM) 1503. The RAM 1503 also stores various programs and data required for the operation of the electronic device 1500. The processing unit 1501, ROM 1502, and RAM 1503 are interconnected via a bus 1504. An input / output (I / O) interface 1505 is also connected to the bus 1504.
[0239] Typically, the following devices can be connected to I / O interface 1505: input devices 1506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 1507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1509. Communication device 1509 allows electronic device 1500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 15 An electronic device 1500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0240] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 1509, or installed from a storage device 1508, or installed from a ROM 1502. When the computer program is executed by a processing device 1501, it performs the functions defined in the methods of embodiments of this disclosure.
[0241] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0242] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0243] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.
[0244] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0245] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0246] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
[0247] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0248] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0249] In a first aspect, according to one or more embodiments of this disclosure, a font size detection method is provided, comprising:
[0250] The image to be detected is processed to obtain the target text content in the image to be detected, and the target rectangular area occupied by the target text content;
[0251] Determine the target calibration font size of the target text content based on the target rectangular region;
[0252] A target correction ratio is determined based on the target text content and the target calibration font size, and the target calibration font size is corrected according to the target correction ratio to obtain the target font size, which is used as the font size detection result of the target text content.
[0253] According to one or more embodiments of this disclosure, determining the target correction ratio based on the target text content and the target calibration font size includes:
[0254] The target text content and the target calibration font size are encoded separately to obtain the target text encoding vector of the target text content and the target font size encoding vector of the target calibration font size;
[0255] The target correction ratio is determined based on the target text encoding vector and the target font size encoding vector.
[0256] According to one or more embodiments of this disclosure, determining the target correction ratio based on the target text encoding vector and the target font size encoding vector includes:
[0257] The target text encoding vector and the target font size encoding vector are fused to obtain a target fused vector, and the target correction ratio is determined based on the target fused vector.
[0258] According to one or more embodiments of this disclosure, determining the target correction ratio based on the target fusion vector includes:
[0259] The target fusion vector is input into a pre-trained font size correction model, and the target correction ratio is output.
[0260] The font size correction model is a model trained on sample images to predict the correction ratio of the font size.
[0261] According to one or more embodiments of this disclosure, the target calibration font number is encoded to obtain a target font number encoding vector for the target calibration font number, including:
[0262] Based on the preset mapping relationship between font size and font size grouping interval, determine the target grouping interval corresponding to the target calibration font size;
[0263] The target grouping interval is encoded to obtain the target font size encoding vector.
[0264] According to one or more embodiments of this disclosure, the target calibration font size is determined based on the target boundary information of the target rectangular region.
[0265] According to one or more embodiments of this disclosure, if the text in the image to be detected is in a horizontal layout, the target boundary information includes the upper boundary information and the lower boundary information of the text in the image to be detected; if the text in the image to be detected is in a vertical layout, the target boundary information includes the left boundary information and the right boundary information of the text in the image to be detected.
[0266] According to one or more embodiments of this disclosure, if the text in the image to be detected is in a horizontal format, the target boundary information is determined by projecting and superimposing the image to be detected in the vertical direction based on the vertical direction of the text in the image to be detected to obtain the target gray value, and then determining the target gray value.
[0267] If the text in the image to be detected is in a vertical format, the target boundary information is determined by projecting and overlaying the image in the horizontal direction of the text in the image to be detected, based on the horizontal direction of the text in the image to be detected, to obtain the target gray value, and then determining the target gray value.
[0268] Secondly, according to one or more embodiments of this disclosure, a text processing method is provided, comprising:
[0269] The sample image is processed to obtain the sample text content in the sample image and the sample rectangular area occupied by the sample text content;
[0270] The sample calibration font size of the sample text content is determined based on the sample rectangular area;
[0271] A font size correction model is trained based on the sample text content and the sample calibrated font size, wherein the font size correction model is used to determine the target correction ratio of the font size of the target text content in the image to be detected.
[0272] According to one or more embodiments of this disclosure, training a font size correction model based on the sample text content and the sample calibrated font size includes:
[0273] The sample text content and the sample calibration font size are encoded separately to obtain the sample text encoding vector of the sample text content and the sample font size encoding vector of the sample calibration font size;
[0274] The sample font size encoding vector and the sample text encoding vector are fused together to obtain a sample fusion vector;
[0275] The font size correction model is generated based on the sample fusion vector and the sample calibration font size.
[0276] According to one or more embodiments of this disclosure, generating the font size correction model based on the sample fusion vector and the sample calibration font size includes:
[0277] The prediction ratio is obtained based on the sample fusion vector, wherein the prediction ratio is used to characterize the font size ratio between the preset standard font size of the predicted sample text content and the font size ratio between the sample calibration font size;
[0278] The true ratio between the preset standard font size and the sample calibration font size is calculated;
[0279] The font size correction model is generated based on the predicted ratio and the actual ratio.
[0280] According to one or more embodiments of this disclosure, generating a sample font size encoding vector for the sample text content based on the sample calibration font size includes:
[0281] Based on the preset mapping relationship between font size and font size grouping interval, determine the sample grouping interval corresponding to the sample calibration font size;
[0282] The sample grouping intervals are encoded to obtain the sample font size encoding vector.
[0283] According to one or more embodiments of this disclosure, the standard font size of the sample is determined based on the sample boundary information of the sample rectangular region.
[0284] According to one or more embodiments of this disclosure, if the text in the sample image is in a horizontal layout, the sample boundary information includes the upper boundary information and the lower boundary information of the text in the sample image; if the text in the sample image is in a vertical layout, the sample boundary information includes the left boundary information and the right boundary information of the text in the sample image.
[0285] According to one or more embodiments of this disclosure, if the text in the sample image is in a horizontal layout, the sample boundary information is determined by projecting and superimposing the sample image in the vertical direction based on the vertical direction of the text in the sample image to obtain the sample gray value, and then determining the gray value based on the sample gray value.
[0286] If the text in the sample image is in a vertical format, then the sample boundary information is determined by projecting and overlaying the sample image horizontally based on the horizontal direction of the text in the sample image to obtain the sample grayscale value, and then determining the grayscale value based on the sample grayscale value.
[0287] Thirdly, according to one or more embodiments of this disclosure, a font size detection device is provided, comprising:
[0288] The first recognition unit is used to perform recognition processing on the image to be detected to obtain the target text content in the image to be detected and the target rectangular area occupied by the target text content.
[0289] The first determining unit is used to determine the target calibration font size of the target text content based on the target rectangular region;
[0290] The second determining unit is used to determine the target correction ratio based on the target text content and the target calibration font size;
[0291] The correction unit is used to correct the target calibration font size according to the target correction ratio to obtain the target font size, which is used as the font size detection result of the target text content.
[0292] According to one or more embodiments of this disclosure, the second determining unit includes:
[0293] The first encoding subunit is used to encode the target text content and the target calibration font size respectively to obtain the target text encoding vector of the target text content and the target font size encoding vector of the target calibration font size;
[0294] A subunit is defined to determine the target correction ratio based on the target text encoding vector and the target font size encoding vector.
[0295] According to one or more embodiments of this disclosure, the determining subunit includes:
[0296] The fusion module is used to fuse the target text encoding vector and the target font size encoding vector to obtain a target fused vector;
[0297] The first determining module is used to determine the target correction ratio based on the target fusion vector.
[0298] According to one or more embodiments of this disclosure, the first determining module is configured to input the target fusion vector into a pre-trained font size correction model and output the target correction ratio;
[0299] The font size correction model is a model trained on sample images to predict the correction ratio of the font size.
[0300] According to one or more embodiments of this disclosure, the first encoding subunit includes:
[0301] The second determining module is used to determine the target grouping interval corresponding to the target calibration font size according to the preset mapping relationship between font size and font size grouping interval;
[0302] The encoding module is used to encode the target grouping interval to obtain the target font size encoding vector.
[0303] According to one or more embodiments of this disclosure, the target calibration font size is determined based on the target boundary information of the target rectangular region.
[0304] According to one or more embodiments of this disclosure, if the text in the image to be detected is in a horizontal layout, the target boundary information includes the upper boundary information and the lower boundary information of the text in the image to be detected; if the text in the image to be detected is in a vertical layout, the target boundary information includes the left boundary information and the right boundary information of the text in the image to be detected.
[0305] According to one or more embodiments of this disclosure, if the text in the image to be detected is in a horizontal format, the target boundary information is determined by projecting and superimposing the image to be detected in the vertical direction based on the vertical direction of the text in the image to be detected to obtain the target gray value, and then determining the target gray value.
[0306] If the text in the image to be detected is in a vertical format, the target boundary information is determined by projecting and overlaying the image in the horizontal direction of the text in the image to be detected, based on the horizontal direction of the text in the image to be detected, to obtain the target gray value, and then determining the target gray value.
[0307] Fourthly, according to one or more embodiments of this disclosure, a text processing apparatus is provided, comprising:
[0308] The second recognition unit is used to perform recognition processing on the sample image to obtain the sample text content in the sample image and the sample rectangular area occupied by the sample text content.
[0309] The third determining unit is used to determine the sample calibration font size of the sample text content based on the sample rectangular area;
[0310] The training unit is used to train a font size correction model based on the sample text content and the sample calibrated font size, wherein the font size correction model is used to determine the target correction ratio of the font size of the target text content in the image to be detected.
[0311] According to one or more embodiments of this disclosure, the training unit includes:
[0312] The second encoding subunit is used to encode the sample text content and the sample calibration font size respectively to obtain the sample text encoding vector of the sample text content and the sample font size encoding vector of the sample calibration font size;
[0313] The fusion subunit is used to fuse the sample font size encoding vector and the sample text encoding vector to obtain a sample fusion vector;
[0314] A sub-unit is generated to generate the font size correction model based on the sample fusion vector and the sample calibration font size.
[0315] According to one or more embodiments of this disclosure, the generating subunit includes:
[0316] The prediction module is used to predict a prediction ratio based on the sample fusion vector, wherein the prediction ratio is used to characterize the font size ratio between the preset standard font size of the predicted sample text content and the font size ratio between the sample calibration font size;
[0317] The calculation module is used to calculate the true ratio between the preset standard font size and the sample calibration font size;
[0318] The generation module is used to generate the font size correction model based on the predicted ratio and the actual ratio.
[0319] According to one or more embodiments of this disclosure, the generation module includes:
[0320] The determination submodule is used to determine the sample grouping interval corresponding to the sample calibration font size based on the preset mapping relationship between font size and font size grouping interval;
[0321] The encoding submodule is used to encode the sample grouping interval to obtain the sample font size encoding vector.
[0322] According to one or more embodiments of this disclosure, the standard font size of the sample is determined based on the sample boundary information of the sample rectangular region.
[0323] According to one or more embodiments of this disclosure, if the text in the sample image is in a horizontal layout, the sample boundary information includes the upper boundary information and the lower boundary information of the text in the sample image; if the text in the sample image is in a vertical layout, the sample boundary information includes the left boundary information and the right boundary information of the text in the sample image.
[0324] According to one or more embodiments of this disclosure, if the text in the sample image is in a horizontal layout, the sample boundary information is determined by projecting and superimposing the sample image in the vertical direction based on the vertical direction of the text in the sample image to obtain the sample gray value, and then determining the gray value based on the sample gray value.
[0325] If the text in the sample image is in a vertical format, then the sample boundary information is determined by projecting and overlaying the sample image horizontally based on the horizontal direction of the text in the sample image to obtain the sample grayscale value, and then determining the grayscale value based on the sample grayscale value.
[0326] Fifthly, according to one or more embodiments of the present disclosure, an electronic device is provided, comprising: at least one processor and a memory;
[0327] The memory stores computer-executed instructions;
[0328] The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the font size detection method as described in the first aspect and various possible designs of the first aspect; or, causing the at least one processor to perform the text processing method as described in the second aspect and various possible designs of the second aspect.
[0329] Sixthly, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, wherein computer-executable instructions are stored therein, which, when executed by a processor, implement the font size detection method as described in the first aspect and various possible designs of the first aspect; or implement the text processing method as described in the second aspect and various possible designs of the second aspect.
[0330] In a seventh aspect, according to one or more embodiments of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the font size detection method as described in the first aspect and various possible designs of the first aspect; or, implements the text processing method as described in the second aspect and various possible designs of the second aspect.
[0331] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0332] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0333] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A font size detection method, characterized in that, include: The image to be detected is processed to obtain the target text content in the image to be detected, and the target rectangular area occupied by the target text content; Determine the target calibration font size of the target text content based on the target rectangular region; A target correction ratio is determined based on the target text content and the target calibration font size, and the target calibration font size is corrected according to the target correction ratio to obtain the target font size, which is used as the font size detection result of the target text content.
2. The method according to claim 1, characterized in that, The step of determining the target correction ratio based on the target text content and the target calibration font size includes: The target text content and the target calibration font size are encoded separately to obtain the target text encoding vector of the target text content and the target font size encoding vector of the target calibration font size; The target correction ratio is determined based on the target text encoding vector and the target font size encoding vector.
3. The method according to claim 2, characterized in that, Determining the target correction ratio based on the target text encoding vector and the target font size encoding vector includes: The target text encoding vector and the target font size encoding vector are fused to obtain a target fused vector, and the target correction ratio is determined based on the target fused vector.
4. The method according to claim 3, characterized in that, Determining the target correction ratio based on the target fusion vector includes: The target fusion vector is input into a pre-trained font size correction model, and the target correction ratio is output. The font size correction model is a model trained on sample images to predict the correction ratio of the font size.
5. The method according to any one of claims 2-4, characterized in that, The target calibration font is encoded to obtain the target font encoding vector, including: Based on the preset mapping relationship between font size and font size grouping interval, determine the target grouping interval corresponding to the target calibration font size; The target grouping interval is encoded to obtain the target font size encoding vector.
6. The method according to any one of claims 1-4, characterized in that, The target calibration font size is determined based on the target boundary information of the target rectangular region.
7. The method according to claim 6, characterized in that, If the text in the image to be detected is in a horizontal layout, the target boundary information includes the upper and lower boundary information of the text in the image to be detected. If the text in the image to be detected is in a vertical layout, the target boundary information includes the left and right boundary information of the text in the image to be detected.
8. The method according to claim 7, characterized in that, If the text in the image to be detected is in a horizontal format, then the target boundary information is determined by projecting and superimposing the image in the vertical direction of the text in the image to be detected based on the vertical direction of the text in the image to be detected, obtaining the target gray value, and determining the target gray value accordingly. If the text in the image to be detected is in a vertical format, the target boundary information is determined by projecting and overlaying the image in the horizontal direction of the text in the image to be detected, based on the horizontal direction of the text in the image to be detected, to obtain the target gray value, and then determining the target gray value.
9. A text processing method, characterized in that, include: The sample image is processed to obtain the sample text content in the sample image and the sample rectangular area occupied by the sample text content; The sample calibration font size of the sample text content is determined based on the sample rectangular area; A font size correction model is trained based on the sample text content and the sample calibrated font size, wherein the font size correction model is used to determine the target correction ratio of the font size of the target text content in the image to be detected.
10. The method according to claim 9, characterized in that, The step of training a font size correction model based on the sample text content and the sample calibrated font size includes: The sample text content and the sample calibration font size are encoded separately to obtain the sample text encoding vector of the sample text content and the sample font size encoding vector of the sample calibration font size; The sample font size encoding vector and the sample text encoding vector are fused together to obtain a sample fusion vector; The font size correction model is generated based on the sample fusion vector and the sample calibration font size.
11. The method according to claim 10, characterized in that, The step of generating the font size correction model based on the sample fusion vector and the sample calibration font size includes: The prediction ratio is obtained based on the sample fusion vector, wherein the prediction ratio is used to characterize the font size ratio between the preset standard font size of the predicted sample text content and the font size ratio between the sample calibration font size; The true ratio between the preset standard font size and the sample calibration font size is calculated; The font size correction model is generated based on the predicted ratio and the actual ratio.
12. A font size detection device, characterized in that, include: The first recognition unit is used to perform recognition processing on the image to be detected to obtain the target text content in the image to be detected and the target rectangular area occupied by the target text content. The first determining unit is used to determine the target calibration font size of the target text content based on the target rectangular region; The second determining unit is used to determine the target correction ratio based on the target text content and the target calibration font size; The correction unit is used to correct the target calibration font size according to the target correction ratio to obtain the target font size, which is used as the font size detection result of the target text content.
13. A text processing device, characterized in that, include: The second recognition unit is used to perform recognition processing on the sample image to obtain the sample text content in the sample image and the sample rectangular area occupied by the sample text content. The third determining unit is used to determine the sample calibration font size of the sample text content based on the sample rectangular area; The training unit is used to train a font size correction model based on the sample text content and the sample calibrated font size, wherein the font size correction model is used to determine the target correction ratio of the font size of the target text content in the image to be detected.
14. An electronic device, characterized in that, include: At least one processor and memory; The memory stores computer-executed instructions; The at least one processor executes the computer execution instructions stored in the memory, causing the at least one processor to perform the font size detection method as described in any one of claims 1 to 8; or, This causes the at least one processor to perform the text processing method as described in any one of claims 9-11.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the font size detection method as described in any one of claims 1 to 8; or implement the text processing method as described in any one of claims 9 to 11.
16. A computer program product comprising a computer program that, when executed by a processor, implements the font size detection method according to any one of claims 1 to 8; or, when executed by a processor, the computer program implements the text processing method according to any one of claims 9 to 11.
Citation Information
Patent Citations
Serial number identification method and device, equipment, and storage medium
CN107358718A
Document word size identification method and device, computer equipment and storage medium
CN115131803A