Information processing apparatus, information processing method, program
The information processing apparatus accurately determines and extracts text data with line breaks matching human visual perception by using character coordinates, size, and progression direction, addressing the issue of incorrect line breaks in PDF files.
Patent Information
- Application Number
- JP2025042132
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-28
- Estimated Expiration
- 2045-03-17
AI Technical Summary
Existing PDF files lack information on line breaks that match visual recognition, leading to incorrect line breaks in extracted text data, making proofreading difficult due to the high degree of freedom in character arrangement, size, and curvature.
An information processing apparatus determines the range of characters included in visually recognized lines by using coordinate information, character size, deformation, and progression direction, and sets extended lines to connect characters within a single line, enabling accurate line determination and extraction.
Enables extraction of text data with line breaks matching human visual perception, suitable for proofreading, even with curved or overlapping character strings.
Smart Images

Figure 0007713678000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, an information processing method, and a program, and particularly relates to the technical field of character extraction from page data files.
Background Art
[0002] A data file called PDF (Portable Document Format) (hereinafter referred to as "PDF file"), which stores the printed state on paper media, is widely used. And in Patent Document 1 below, a technique for determining the content of character list output according to the presence or absence of character overlap due to the arrangement of text boxes in a PDF file is disclosed.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] By the way, character extraction for extracting text data from a PDF file is a useful technique, for example, when proofreading printed materials. For example, when a PDF file containing images and characters is created as a print layout for product labels, packages, advertising flyers, etc., it is necessary to check whether the content described in the PDF file matches the original manuscript submitted. In this case, character data (text data) is extracted from the PDF file, and proofreading is performed by checking whether the extracted character data matches the submitted text.
[0005] However, in a PDF file, information clearly indicating the separation of strings that are recognized as "lines" by humans in the printed state is not necessarily added. Although PDF files contain various additional information (metadata) regarding character data, information on line breaks that match visual recognition is not always attached. Also, in the case of PDF files intended for printing on product packages, etc., the arrangement, size, overlap, etc. of characters have a high degree of freedom, and the character strings may be curved.
[0006] Due to these circumstances, even if character data is read from a PDF file, the line breaks will be significantly different from those in the original manuscript, resulting in text data that is inconvenient for proofreaders.
[0007] Therefore, an object of the present invention is to enable character extraction with line breaks similar to those of lines when humans view page data such as PDF files.
Means for Solving the Problem
[0008] The information processing apparatus according to the present invention performs continuity determination based on the arrangement positions of each character using at least the coordinate information of the character, the size and deformation information of the character, and the information on the progress direction of the character string for the character data included in the page data file, and determines the range of characters included in the lines visually recognized in the printed state. It also includes an arithmetic unit that performs a second process of separating lines for each range of characters determined in the first process, extracting text data, and outputting it. And in the first process, the calculation unit sets a determination reference line that is a line in the progress direction of the character string within the character frame for each character, sets an extended line obtained by extending one or both of the line width and line length of the determination reference line of each character string, and determines that the characters within the range where the extended lines are continuous are the characters included in one line. For example, based on the continuity of the arrangement positions of each character in a page data file such as PDF, the range of character strings visually recognized as lines after the PDF data is printed is determined.
Effect of the Invention
[0009] According to the present invention, text data can be extracted with line divisions as humans would view them when a page data file is printed on paper media. Therefore, highly suitable text data extraction for proofreading page data files can be realized.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Modes for Carrying Out the Invention
[0011] Hereinafter, an information processing apparatus as an embodiment for performing a process of extracting character data from a PDF file will be described.
[0012] In the present disclosure, a "line" indicates a string range as a sequence of characters aligned in a certain one direction, and any of the horizontal direction, vertical direction, and diagonal direction is regarded as a "line". The "advancing direction" of a character string refers to information on whether the characters are written vertically or horizontally, that is, the direction of the reading order, and the "reverse advancing direction" refers to the direction opposite to the advancing direction (the direction from the end of the character string to the beginning, that is, the direction of going back against the reading order).
[0013] For a normal Japanese character string, in the case of horizontal writing, a row of horizontally aligned characters is "one line", and the advancing direction is the direction from left to right. In the case of vertical writing, a column of vertically aligned characters is "one line", and the advancing direction is the direction from top to bottom. However, the PDF file targeted in this embodiment is a PDF file including images and characters as a print layout such as product labels, packages, and advertising flyers, and the advancing direction (reading order direction) may be diagonal, curved, from right to left, or from bottom to top.
[0014] The character deformation information refers to information corresponding to parameters when the characters are deformed, such as character enlargement, rotation, and inclination.
[0015] <Configuration of Information Processing Apparatus> FIG. 1 shows a configuration example of the information processing apparatus 1. The information processing apparatus 1 is a device capable of information processing, such as a computer device. Specifically, the configuration of the information processing apparatus 1 is also assumed to be a workstation, a personal computer, a mobile terminal device such as a smartphone or a tablet, etc. Further, the information processing apparatus 1 may be a computer device configured as a server device or an arithmetic device in cloud computing.
[0016] The information processing apparatus 1 includes an arithmetic unit 2, a ROM (Read Only Memory) 3, a RAM (Random Access Memory) 4, and a non-volatile memory 5.
[0017] The arithmetic unit 2 is composed of a processor such as a CPU (Central Processing Unit), for example. For example, the arithmetic unit 2 executes various control processes and arithmetic processes according to programs stored in the ROM 3 or the non-volatile memory 5 such as an EEP-ROM (Electrically Erasable Programmable Read-Only Memory), or programs loaded from the storage unit 8 to the RAM 4. The RAM 4 also appropriately stores data and the like necessary for the arithmetic unit 2 to execute various processes.
[0018] The information processing apparatus 1 may include an input unit 6, a display unit 7, a storage unit 8, a communication unit 9, and a recording / reproducing unit 10. For example, as the input unit 6, various operators and operation devices such as a keyboard, a mouse, a touch panel, a touch pad, and a remote controller are assumed. An operation of the user is detected by the input unit 6, and a signal corresponding to the input operation is interpreted by the arithmetic unit 2. The character extraction process described later is executed in response to the user designating a PDF file to be processed and instructing the character extraction process by an operation using the input unit 6.
[0019] The display unit 7 is formed of, for example, an LCD (Liquid Crystal Display) or an organic EL (electro-luminescence) panel and performs various displays. The display unit 7 is composed of, for example, a display device provided on the housing of the information processing apparatus 1 or a separate display device connected to the information processing apparatus 1. On the display unit 7, for example, display of a PDF file that is the target of the character extraction process and display of text data as the extraction process result are performed based on the control of the arithmetic unit 2.
[0020] The storage unit 8 indicates a relatively large-capacity storage device composed of an HDD (Hard Disk Drive), an SSD (Solid State Drive), or the like. The storage unit 8 can store various data and programs. A database can also be configured in the storage unit 8. The communication unit 9 performs communication processing via a transmission path such as the Internet, and communication such as wired / wireless communication and bus communication with various devices such as an external database, an editing device, and an information processing device. In the storage unit 8 and the communication unit 9, storage and transmission / reception of PDF files to be processed in the present embodiment, extracted text data, and the like are performed.
[0021] The recording / reproducing unit 10 is a media drive device that performs recording and reproduction on a removable medium 11 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory. With the recording / reproducing unit 10, various data files and various computer programs can be read from the removable medium 11. The read data is stored in the storage unit 8, or characters and images included in the data are output by the display unit 7. Also, computer programs and the like read from the removable medium 11 are installed in the storage unit 8 as necessary.
[0022] In this information processing apparatus 1, for example, a software program for the processing of the present embodiment, that is, the character extraction processing from a PDF file described later, can be installed via network communication by the communication unit 9 or via the removable medium 11. Alternatively, the software program may be stored in advance in the ROM 3, the storage unit 8, or the like.
[0023] <Comparison Example of PDF File and Character Extraction> FIG. 2A shows an example of a PDF file created so that a label, a package, or the like for a container of a certain product is generated by printing.
[0024] In this PDF file, the character data 22 and the image data 21 are arranged as in the actual printed matter. As the character data 22, for example, as shown in FIG. 3A, various characters such as the type of product and the characters of ingredients are described over a plurality of lines. Note that this PDF file is a file for producing the surface of a product container which is a curved surface, and the character string is taken as an example that is slightly curved.
[0025] In the proofreading operation of the PDF file, it is confirmed whether the characters described as in FIG. 3A match the characters described in the original manuscript. Therefore, a character extraction process is performed on the PDF file of FIG. 2A, and a process of extracting only the character data 22 is performed. However, if the extraction process is simply performed, for example, text data as shown in FIG. 3B will be extracted and displayed.
[0026] In FIG. 3B, for convenience of explanation, line numbers are attached. Only the text part is actually extracted. This FIG. 3B shows an example where the character string does not have the same line break as that in FIG. 3A. For example, while the first line in FIG. 3A is "Type: Fermented milk, non-fat milk solids", the extracted text in FIG. 3B is divided into three lines from the beginning: "Type: Fermented", "Milk, non-fat milk solid", and "Form". This is due to the fact that the character data 22 in FIG. 3A is a curved character string and the line break information is not accurately added. As described above, the PDF file only stores the arrangement of characters and images in the printed state, and does not accurately hold the line information as additional data. Also, the character strings are not always arranged in a straight line, and the sizes, fonts, etc. are diverse. Therefore, even if a character data extraction process is performed, the range that appears as one line visually is not necessarily output as one line of text data.
[0027] And the text data output as in FIG. 3B is not suitable for the proofreading operation. Due to the different line breaks, it is difficult to collate with the original manuscript. Ideally, as shown in FIG. 4, it is desirable to be able to output text data in which the character string and the line break in FIG. 3A match. Therefore, in this embodiment, as shown in FIG. 4, text data can be extracted from a PDF file with the same line breaks as in the printed state.
[0028] Also, the order of character strings in character data in a PDF file is diverse. FIG. 5 shows an example of the order of character strings. Here, examples of 0 degrees, 180 degrees, 90 degrees, 270 degrees, and arbitrary degrees of character rotation in the case of horizontal writing are shown respectively, and examples of 0 degrees, 180 degrees, 90 degrees, 270 degrees, and arbitrary degrees of character rotation in the case of vertical writing are shown respectively. In the character extraction process of the embodiment, even for such diverse orders of character strings, appropriate line determination is performed to realize text data extraction in accordance with the reading order direction.
[0029] <Character Extraction Process of the Embodiment> FIG. 6 shows an example of a character extraction process by the information processing apparatus 1 of the embodiment. The process in FIG. 6 is an example of the process executed by the calculation unit 2 on the PDF data targeted for processing in response to a user operation. The calculation unit 2 performs the process in FIG. 6 based on a software program for character extraction.
[0030] In step S101, the calculation unit 2 acquires a PDF file to be the target of the character extraction process according to the user's operation. In step S102, the calculation unit 2 acquires additional information from the PDF file targeted for processing. Specifically, the calculation unit 2 acquires coordinate values (character coordinates (x, y) indicating the center of the character) for each character as character data 22, deformation information (transform) such as character enlargement / rotation / slant, information (trim) on the margins around the character and figure, information (size) on the character size, information (adv) on the character width ratio with respect to the character size, information (wmode) on the progress direction such as vertical or horizontal writing, etc.
[0031] In step S103, the arithmetic unit 2 sets a character frame 23 for each character data 22 on a character-by-character basis. FIG. 7 shows a state in which a character frame 23 is set for each character data 22. To avoid complicating the figure, only some characters and character frames are labeled, but a character frame 23 is set for all character data 22 to be read.
[0032] This setting of the character frame 23 is actually a process of setting character contour coordinates, which are the coordinate values of the four corners of the frame. For example, for the character data 22 of the character "あ" in FIG. 8, as the character frame 23, a process of obtaining the coordinates of four points: ll (lower left), ul (upper left), ur (upper right), and lr (lower right) is involved. The arithmetic unit 2 calculates these coordinate values using the above-mentioned character coordinates (x, y), character deformation information, information on the margins around the character and the figure, character size information, and information on the character width ratio with respect to the character size. Thereby, regardless of the size, rotation, surrounding margins, etc. of the character data 22, the character contour coordinates can be calculated with high precision. Note that as long as at least the character coordinates, character deformation information, and size information can be obtained, the character contour coordinate values can be calculated with sufficient accuracy. If there is margin information or character width ratio information, the accuracy of the character contour coordinates can be further enhanced.
[0033] In step S104, the arithmetic unit 2 sets the extraction range of the character data 22. This is a process of, for example, filling in each character frame 23 set as shown in FIG. 7 to set the area for character extraction. By performing the process of filling in the character frames 23 of all character data 22 described in the PDF file, the area for character extraction can be shown on the PDF file as the extraction area 24 with diagonal lines in FIG. 2B. For example, by displaying the state of FIG. 2B on the display unit 7, the area for character extraction can be presented to the user.
[0034] In step S105, the arithmetic unit 2 sets the determination reference line 30. FIG. 9 shows the setting of the determination reference line 30. The determination reference line 30 is a line used for the continuity determination based on the arrangement positions of the respective characters. The arithmetic unit 2 sets the determination reference line 30 for each character based on the character frame 23 (character contour coordinates) and information on the direction of progress such as whether it is vertical writing or horizontal writing.
[0035] FIG. 9A is an example when it is determined that the character string is horizontal writing from left to right based on the information on the direction of progress. In this case, the arithmetic unit 2 sets the determination reference line 30 with the midpoint of the coordinates ul and ll as the start point coordinate SP and the midpoint of the coordinates ur and lr as the end point coordinate EP. FIG. 9B is an example when it is determined that the character string is vertical writing from top to bottom based on the information on the direction of progress. In this case, the arithmetic unit 2 sets the determination reference line 30 with the midpoint of the coordinates ul and ur as the start point coordinate SP and the midpoint of the coordinates ll and lr as the end point coordinate EP. As a result, the determination reference line 30, which is the line in the direction of progress of the character string within the character frame 23 of each character, can be set. For the calculation of these determination reference lines 30 (calculation of the start point coordinate SP and the end point coordinate EP), the coordinates of the character frame 23, the information on the direction of progress (wmode), and the inclination (direction x, direction y) = (cosθ, sinθ) in the character deformation information are used.
[0036] Note that the determination reference line 30 is not limited to a line passing through the center of the character frame 23 as shown in FIGS. 9A and 9B. As shown in FIGS. 9C and 9D, it may pass through a line on the character frame 23. For example, FIG. 9C is a case when it is determined that the character string is horizontal writing from left to right based on the information on the direction of progress, and the arithmetic unit 2 may set the determination reference line 30 with the coordinate ll as the start point coordinate SP and the coordinate lr as the end point coordinate EP. Although not shown, the determination reference line 30 may be set with the coordinate ul as the start point coordinate SP and the coordinate ur as the end point coordinate EP.
[0037] Also, FIG. 9D shows a case where it is determined that the character string is in a vertical writing style from top to bottom based on the information on the direction of progress. The arithmetic unit 2 may set the determination reference line 30 with the coordinate ur as the start point coordinate SP and the coordinate lr as the end point coordinate EP. Although not shown, the determination reference line 30 may be set with the coordinate ul as the start point coordinate SP and the coordinate ll as the end point coordinate EP.
[0038] Even in other cases, within the character frame 23, a line corresponding to the direction of progress of the character string may be set as the determination reference line 30 using the coordinates on the frame line of the character frame 23.
[0039] Subsequently, the arithmetic unit 2 sets an extended line obtained by extending the line length and line width of the determination reference line 30 in step S106. This is a process of extending the determination reference line 30 set within the character frame 23 of each character by a predetermined length or thickening the line width to connect them, so that the characters within the range where the determination reference lines are continuous can be determined as characters within the range included in one line visually.
[0040] The PDF file also contains information on connected characters, but this is not information on the "line" recognized visually. Even if it appears to be one line visually, it is not necessarily the case that one line's worth of characters is connected, and there may be multiple connected ranges. This is illustrated in FIG. 10. Assume a case where the character string "Haru wa akebono" shown in FIG. 10A is regarded as one line visually. However, on the PDF file, assume that it is composed of a connected part "Haru wa" and a connected part "akebono" as shown in FIG. 10B. In order to recognize these multiple connected parts as one line, the arithmetic unit 2 extends the determination reference line 30.
[0041] First, as shown in FIG. 10C, the arithmetic unit 2 obtains the center coordinates CP between the characters within the line from the above-mentioned coordinates (ul, ll, ur, lr) of the character frame 23 and the center coordinates (x, y) of the character data 22. Then, as shown in FIG. 10D, the center coordinates CP are connected. The above can be said to be the result of extending and connecting the determination reference lines 30 of individual characters within the range of the connection information between the lines.
[0042] However, just by this alone, there may be cases where not all ranges of a single line can be concatenated. As shown in FIG. 10B, if "Spring" and "dawn" are in different concatenated parts, there will be a gap (the dashed circle part) between the determination reference lines 30 concatenated in their respective ranges as shown in FIG. 10E. To connect such a gap, the determination reference line 30 is extended in the reading direction of the character string by, for example, about 12.5% of the character size. Or, the line width (thickness of the line) is thickened. For example, a 1-pixel-sized line is made into a 3-pixel-sized line. The line extended in this way is defined as the extended line 31. FIG. 10F shows the extended line 31 (the thick line in the hatched area) obtained by extending the determination reference line 30. This extended line 31 is a line that concatenates the ranges of characters in a single line. And characters that are partially included in the determination reference line 30 within the contour range of the extended line 31 are determined to be characters included in the range of a single line.
[0043] FIG. 10 shows an example where the character string is advancing horizontally. The same applies even when the character string advances in a curved manner. This is illustrated in FIG. 11. For example, for the character string as shown in FIG. 7 above, as shown in FIG. 11A, the determination reference line 30 is extended to connect the center coordinates CP between characters. This enables the connection of the determination reference line 30 even if the positions of the characters are shifted. However, as shown in FIG. 11B, within the visually recognized range of a single line, there may be a non-connectable empty part (the dashed circle part). Therefore, it is extended by a length of a predetermined percentage (p%) of the character size, and also, for example, the line width is thickened by a number of extended pixels of a predetermined percentage (q%) of the character size. Let p and q be values set for extension and enlargement.
[0044] FIG. 11C shows the extended line 31 (the hatched area) obtained by extending the determination reference line 30 in the line direction and the width direction. By setting the extended line 31 in this way, characters that are partially included in the determination reference line 30 within the contour range of the extended line 31 can be determined to be characters included in the range of a single line. In particular, for the setting of the extended line 31, the extension of the line length and the line width of the determination reference line 30 with an expansion rate according to the character size is suitable for connecting each character in the range of a single line with the extended line 31.
[0045] Figure 12 shows the state where after setting the character frame 23 as in Figure 7, the processing up to the setting of the extension line 31 in step S106 has been performed. In step S107 of Figure 6, the arithmetic unit 2 determines the characters belonging to one line in the state where the extension line 31 is set as described above. That is, the characters whose determination reference line 30 is included in the range of the extension line 31 are regarded as the characters included in one line.
[0046] Specifically, for example, by comparing the coordinate range (starting point coordinate to ending point coordinate) of each character's determination reference line 30 with the coordinate range of the area as the extension line 31, it is possible to determine whether the determination reference line 30 is included.
[0047] Then in step S108, the arithmetic unit 2 extracts characters based on the determination of the range of each line and outputs text data. As a result, as shown in Figure 4, text data that is line - broken in the same way as the visual line - break can be obtained. In particular, by taking the range connected by the extension line 31 as one line, even if there are steps in the positions of the characters of the character string or the character string is curved, the line in the reading order direction of the characters can be determined.
[0048] Note that without extending between the characters of the determination reference line 30 as in Figure 10D, by setting the extension line 31 by expanding the determination reference line 30 in the line direction and width direction within the character frame 23, the characters can be connected.
[0049] Here, various cases regarding the determination of the characters in the range of one line by the determination reference line 30 and the extension line 31 will be further explained. Figure 13 shows an example of determining characters that visually appear as one line while being divided into a plurality of text boxes on a PDF file.
[0050] For example, assume that the character string "5 minutes 30 seconds to 6 minutes" in FIG. 13A is divided into text boxes TB1, TB2, TB3, and TB4 with the texts "5 30", "minutes seconds to", "6", and "minutes" respectively, and they are superimposed to form the string. In contrast, for each individual character, the setting of the character frame 23, the setting of the determination reference line 30, connection, and the setting of the extension line 31 are performed in the same manner as described above. Then, the determination reference line 30 of each character is included within the range of the contour of the continuous extension line 31 within the range of one line (FIGS. 13B and 13C). As a result, as shown in FIG. 13D, text data for one line can be output. That is, in the algorithm using the extension line 31 that extends the determination reference line 30, even if text boxes with different character sizes, positions, etc. are superimposed, it is possible to handle the determination of one visually recognized line.
[0051] Also, assume that FIG. 14A is divided into a text box TB11 containing only numbers and a text box TB12 containing only units. By setting the extension line 31 as shown in FIG. 14B for each row and determining the determination reference line 30 included in the extension line 31 as shown in FIG. 14C, the range of one line can be determined. As a result, as shown in FIG. 14D, text data reflecting the line breaks in terms of visual recognition for each row can be output. That is, in the algorithm using the extension line 31 that extends the determination reference line 30, even when multiple rows of text boxes are arranged in the row direction, it is possible to handle the determination of one visually recognized line.
[0052] FIG. 15 shows the processing corresponding to a tabular character string. In the case of tabular character data as shown in FIG. 15A, for example, there is a relatively long blank between an item part such as "Energy" and a numerical part such as "266 kcal", and there may be a case where the extension line 31 is not continuous even when the determination reference line 30 is extended by a predetermined amount. FIG. 15 shows an example where the extension line 31 is interrupted between the item part and the numerical part. In such a case, the extension line 31 is further extended. For example, in the process within step S106 of FIG. 6, the arithmetic unit 2 sets a condition of "a character string that starts with a numerical value and ends with a unit", and in the case of a match, performs a process of extending the extension line 31 of the character string in the reverse direction of progress. As a result, a connection portion 31a is formed as shown in FIG. 15C, and the range of one line is continued by the extension line 31. As a result, as shown in FIG. 15D, even for a tabular character string, it becomes possible to output text data reflecting each visible line.
[0053] FIG. 16 shows text data extraction from a largely curved curved character string. For a spiral character string as shown in FIG. 16A, in the process of FIG. 6, the setting of the character frame 23, the setting of the determination reference line 30, the connection, and the setting of the extension line 31 are performed. Then, the determination reference line 30 of each character is included within the range of the contour of the extension line 31 that is continuous within the range of one line (FIGS. 16B and 16C). As a result, as shown in FIG. 16D, it is possible to output text data of one line. That is, in the algorithm using the extension line 31 that extends the determination reference line 30, even if the progress direction of the character string is largely curved, it is possible to cope with the determination of one visible line.
[0054] <Effects of the Embodiment> According to the information processing apparatus of the above embodiment or the information processing method performed by the information processing apparatus as described above, the following effects can be obtained.
[0055] The arithmetic unit 2 of the information processing apparatus 1 of the embodiment performs a first process (steps S102 to S107 in FIG. 6) of determining the continuity based on the arrangement positions of each character using at least the coordinate information of the character (character coordinates (x, y)), the size of the character (size), the deformation information (transform), and the information on the progress direction of the character string (wmode) for the character data included in the PDF file, and determining the range of characters included in the line visually recognized in the printed state. Further, a second process (step S108) of dividing the lines for each range of characters determined in the first process and extracting and outputting the text data is performed. As a result, regardless of line breaks, string curvature, or text box overlap in the character string of the PDF file, text data can be extracted in line divisions as visually recognized when the PDF file is printed on paper, packaging materials, etc. In particular, by using the coordinate information of characters, the rotation information of characters, and the information on the progression direction of character strings, regardless of the direction in which the character string is formed, the "line" as the direction in which the viewer reads can be appropriately determined, enabling text data output that conforms to the visually recognized lines. Therefore, highly suitable text data extraction for the purpose of proofreading the text described in the PDF file by comparing it with the submitted text can be realized.
[0056] In the embodiment, in the first process, the arithmetic unit 2 sets a reference line 30 that is a line in the progression direction of the character string within the character frame 23 of each character for each character, and sets an extended line 31 obtained by extending one or both of the line width and line length of the reference line 30 of each character string. An example was given in which characters within the range where the extended lines 31 are continuous are determined to be characters included in one line. The characters within the range where the extended lines 31 are continuous are defined as characters at least a part of whose reference line 30 is included within the contour of the extended line 31. By setting the extended line 31 obtained by extending one or both of the line width and line length of the reference line 30 drawn in the character string direction, even when the character string is curved, has steps, or multiple text boxes overlap, the visually recognized lines can be determined based on the continuity of the extended lines 31. As a result, even if the character string is non-linear, different fonts are mixed, or the arrangement is irregular, appropriate text data output corresponding to the visually recognized lines becomes possible.
[0057] In the embodiment, in the first process, when the start of the character string continuous with the extended line 31 is a number and the end is a character indicating a unit, an example was given in which the extended line 31 of the character string is extended in the reverse direction of the progression of the character string and made continuous with the extended line 31 existing in the reverse direction of the progression (see step S106 in FIG. 6 and FIG. 15). For example, in the case of tabular data, there is usually a wide space between the item name and the numerical value. Therefore, for a character string that starts with a number and ends with a unit, the extension line 31 is extended forward in the reverse direction of the character string progress and connected to the character string of the item name. As a result, for tabular character data as well, text data output can be appropriately performed as one line.
[0058] In the embodiment, an example was given in which the extension amount of the determination reference line 30 when setting the extension line 31 is set according to the character size of the target character string. That is, depending on the size of the character data recorded in the PDF, how much the determination reference line 30 is extended in the width direction or the line direction is set according to the character size. As a result, regardless of the size of the character size, the extension line 31 for line determination can be appropriately set. Note that it is also conceivable to set the extension amount as a fixed value.
[0059] The program of the embodiment is a program that causes the arithmetic unit 2 of the information processing apparatus 1 to execute processing as shown in FIG. 6. By such a program, the above-described information processing apparatus 1 can be realized by various computer apparatuses. For example, a personal computer, a portable terminal device such as a smartphone or a tablet, etc. can function as the information processing apparatus 1 of the present disclosure.
[0060] The program can be pre-recorded in a recording medium built in a device such as a computer device or in a ROM in a microcomputer having a CPU. Further, such a program can be stored (recorded) temporarily or permanently in various removable recording media. Further, such a program can be installed from a removable recording medium to a personal computer or the like, or can also be downloaded from a download site via a network such as a LAN (Local Area Network) or the Internet.
Explanation of Signs
[0061] 1 Information processing apparatus 2 Arithmetic unit 20 PDF data 21 Image data 22 Character data 23 Character frame 24 Extraction area 30 Judgment reference line 31 Extension line
Claims
1. For the character data included in the page data file, perform a continuity determination based on the arrangement position of each character using at least the coordinate information of the character, the size and deformation information of the character, and the information on the progress direction of the character string, and determine the range of characters included in the line visually recognized in the printed state; a first process; A second process of dividing lines for each range of characters determined in the first process, extracting text data, and outputting it; An arithmetic unit that performs the above; In the first process, the arithmetic unit For each character, set a reference line that is a line in the progress direction of the character string within the character frame of each character, set an extended line obtained by extending one or both of the line width and line length of the reference line of each character string, and determine that the characters within the range where the extended lines are continuous are the characters included in one line An information processing device.
2. In the first process, the arithmetic unit When the head of the character string continuous with the extended line is a number and the tail is a character indicating a unit, perform a process of extending the extended line of the character string in the reverse direction of the progress of the character string and making it continuous with the extended line existing in the reverse direction of the progress The information processing device according to Claim 1.
3. The extension amount of the reference line when setting the extended line is set according to the character size of the target character string The information processing device according to Claim 1.
4. An information processing device For the character data included in the page data file, perform a continuity determination based on the arrangement position of each character using at least the coordinate information of the character, the size and deformation information of the character, and the information on the progress direction of the character string, and determine the range of characters included in the line visually recognized in the printed state; a first process; A second process of dividing lines for each range of characters determined in the first process, extracting text data, and outputting it; Perform the above, In the first process, For each character, set a reference line that is a line in the progress direction of the character string within the character frame of each character, set an extended line obtained by extending one or both of the line width and line length of the reference line of each character string, and determine that the characters within the range where the extended lines are continuous are the characters included in one line An information processing method.
5. For the character data included in the page data file, perform a continuity determination based on the arrangement position of each character using at least the coordinate information of the character, the size and deformation information of the character, and the information on the progress direction of the character string, and determine the range of characters included in the line visually recognized in the printed state; a first process; Causing the information processing apparatus to execute a second process of dividing lines for each range of characters determined in the first process, extracting text data, and outputting the text data. In the first process, for each character, setting a determination reference line that is a line in the advancing direction of the character string within the character frame of each character, setting an extended line obtained by extending one or both of the line width and line length of the determination reference line of each character string, and causing the information processing apparatus to execute a process of determining that characters in a range where the extended lines are continuous are characters included in one line. Program.
Citation Information
Patent Citations
Image rearrangement method, image rearrangement system and image rearrangement program
JP2014085689A
Program for document image processing and image processor and character recognition device using the program
JP2016170677A
Sentence extraction apparatus and program
JP2021163159A