Document information processing device and program
The document information processing apparatus addresses the inefficiencies in processing layout-processed documents by using a generative AI model to estimate and translate text data, thereby reducing manual work and improving translation accuracy.
Patent Information
- Application Number
- JP2023200719
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-28
- Publication Date
- 2025-06-09
AI Technical Summary
Existing document information processing systems face challenges in efficiently processing documents that have undergone layout processing, such as PDF files, which require manual operations to specify writing direction and arrangement order, leading to increased manual work and reduced processing efficiency.
A document information processing apparatus and program that includes a receiving unit for receiving document information with layout-processed character data, an estimating unit for estimating text data based on the character data, and a processing unit for subjecting the estimated text data to translation processing, utilizing generative AI models to automate the extraction and processing of text data.
The solution significantly reduces manual work and improves processing efficiency by automating the extraction and processing of text data from layout-processed documents, ensuring accurate translation and reducing the need for post-editing.
Smart Images

Figure 2025086623000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a document information processing apparatus and a program for processing document information to be used for translation and the like.
Background Art
[0002] In recent years, translation support tools widely used in the translation industry construct a workflow between external contractors such as translators and proofreaders under the instructions of project managers of translation companies to manage the translation process. In such a translation support tool, when receiving a customer's manuscript from a project manager, it performs segmentation processing (processing of dividing into sentences or clauses) on the manuscript and provides it to translators and the like.
[0003] On the other hand, in recent years, the accuracy of machine translation has improved, and attempts have been made to directly use the segmentation processing results by such translation support tools as the input of machine translation software.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, the target of segmentation processing, such as a customer's manuscript in translation, may be something that has undergone layout processing. Specifically, it may be something in which character data has been layout-processed, such as PDF (Portable Document Format). When layout processing has been performed, encoding of character data, adjustment of fonts, arrangement, and placement are appropriately carried out, so there is a problem that artificial operations are often required in specifying the writing direction and arrangement order.
[0006] The present invention has been made in view of the above circumstances, and one of its objects is to provide a document information processing apparatus and a program that can reduce manual work and improve processing efficiency.
Means for Solving the Problems
[0007] One aspect of the present invention for solving the problems of the above conventional example is a document information processing apparatus, including: a receiving means for receiving document information including character data subjected to predetermined layout processing; an estimating means for estimating text data to be extracted from the received document information based on the character data included in the document information; and a processing means for subjecting the estimated text data to predetermined translation processing.
Effects of the Invention
[0008] According to the present invention, processing efficiency can be improved.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Embodiments for Carrying Out the Invention
[0010] Embodiments of the present invention will be described with reference to the drawings. A document information processing apparatus 1 according to an example of an embodiment of the present invention is realized using a general computer apparatus including a control unit 11, a storage unit 12, an operation unit 13, a display unit 14, and a communication unit 15, as illustrated in FIG. 1.
[0011] The control unit 11 includes a program control device such as a processor and operates according to a program stored in the storage unit 12. The control unit 11 of the present embodiment receives document information including character data subjected to predetermined layout processing by processing according to this program, and estimates text data to be extracted from the received document information based on the character data included in the received document information. In an example of the present embodiment, this control unit 11 may subject the estimated text data to a predetermined translation process. The detailed operation of this control unit 11 will be described later.
[0012] The storage unit 12 includes a memory device and a disk device and holds a program executed by the control unit 11. The operation unit 13 is a keyboard, a mouse, etc., receives a user's operation, and outputs information representing the content of the operation to the control unit 11.
[0013] The display unit 14 is a display or the like and outputs information for display according to an instruction input from the control unit 11. The communication unit 15 is a network interface or the like and outputs information received via the network to the control unit 11. Further, this communication unit 15 transmits information via the network according to an instruction input from the control unit 11.
[0014] Next, the operation of the control unit 11 of the present embodiment will be described. The control unit 11 of the present embodiment realizes a configuration including a reception unit 21, a character data extraction unit 22, an estimation processing unit 23, and an output unit 24 functionally by operating according to a program stored in the storage unit 12.
[0015] Here, the receiving unit 21 receives document information including character data subjected to predetermined layout processing. As a specific example, this receiving unit 21 receives PDF (Portable Document Format) data including character data as document information including character data subjected to predetermined layout processing. This PDF is a file format defined by Adobe in the United States. In order to be oriented towards final output or to emphasize portability, layout processing such as specifying the coordinates within the drawing space for character data to arrange characters and utilizing ligatures to improve readability is performed on the character data. Therefore, for example, when converting original data to PDF, the context (such as the order of character arrangement) of the original data may be lost.
[0016] The character data extraction unit 22 extracts the character data included in the document information (such as PDF data) received by the receiving unit 21. Such extraction of character data can be performed using widely known software such as apache tika. The character data extraction unit 22 further divides the extracted character data into segments that are a predetermined unit (such as sentence unit, etc.) by a process conforming to SRX (Segmentation Rules eXchange), and outputs it as a file in a predetermined format.
[0017] Here, the process performed by the character data extraction unit 22 may be, for example, machine translation (the same process as the machine translation to be performed later, but to distinguish it from the machine translation performed on the output of the estimation processing unit 23, when performing translation processing here, the translation processing here is called preliminary translation processing). In this case, the character data extraction unit 22 outputs the results of the preliminary translation processing for each segment in the form of xliff (XML Localization Interchange Format).
[0018] When the document information received by the receiving unit 21 is document information in PDF format or the like, the character data extraction unit 22 extracts character data from this document information. In this case, the character data may include ligatures or combining characters in Unicode (such as those in which Japanese hiragana and dakuten are represented by two characters). The character data extraction unit 22 uses the character data obtained by directly extracting these characters as the target (original data) of the preliminary translation process. That is, when the character data extracted from the document information includes a ligature (in the following example, a ligature in which f and i are combined), the original data (the content of the source tag) in xliff format output by the character data extraction unit 22 may include the following (however, the following fi part is a ligature): signi<0001>fi<!--0001-->cant
[0019] This example should originally be the text data "significant", but since "fi" is treated as a ligature, the above text data is what has been extracted.
[0020] The estimation processing unit 23 estimates the text data that should have been originally extracted from the document information received by the receiving unit 21 based on the character data extracted by the character data extraction unit 22 (the data before the above-mentioned predetermined processing, or the original data if the character data extraction unit 22 performs preliminary translation processing).
[0021] As an example, based on instructions in the language, this estimation processing unit 23 uses a generation AI model in a state machine-learned to generate sentences, inputs the character data extracted by the character data extraction unit 22 into the generation AI model, and instructs it to estimate the text data to be extracted from the document information. Since the method of estimation processing using such a generation AI model can adopt a widely known processing method that utilizes machine learning results (so-called artificial intelligence), detailed explanation here is omitted.
[0022] Here, the generative AI model may be, for example, a machine learning model using a large language model (such as a Transformer model that has learned a large amount of text data). Such generative AI models include, for example, generative AI models that are large language models based on Generative Pre-trained Transformer, such as ChatGPT (trademark) of OpenAI in the United States. In this embodiment, these generative AI models may be processed by being called via an API (Application Program Interface).
[0023] When, for example, the character data extraction unit 22 performs pre-translation processing and outputs a file in the xliff format, the estimation processing unit 23 extracts the original text data (the part with the source tag) included in the xliff file and instructs the generative AI model to remove the tagged part from the extracted character data. This instruction (prompt) is, for example, "In the following English sentences, please correct English sentences that are concatenated before and after without tags and should only output the corrected sentence." and so on.
[0024] Note that the content in parentheses shows the translation of the instructions when the generative AI model accepts instructions in English. Also, for the part marked as "English text" in the prompt, it may vary depending on the language of the sentence represented by the input character data (or the text data to be extracted). That is, for example, based on the language of the sentence represented by the input character data (or the text data to be extracted), when the character data extracted by the character data extraction unit 22 contains tagged parts, the estimation processing unit 23 instructs the generative AI model to remove the tagged parts, and inputs the character data extracted by the character data extraction unit 22 to the generative AI model.
[0025] Then, as the output of the estimation processing unit 23, it acquires and outputs the text data that is estimated to be the text data to be extracted from the document information. As a result, the estimation processing unit 23 outputs the result of estimating the text data in which the ligatures or combined character strings included in the character data are replaced with the characters of a predetermined character set.
[0026] The output unit 24 outputs the output of the estimation processing unit 23 to a predetermined output destination. This output destination may be, for example, a file. Also, the output unit 24 may cause translation processing to be performed by sending the output of the estimation processing unit 23 to a machine translation system using a machine learning model via the communication unit 15.
[0027] This embodiment basically has the above configuration. However, in document information including character data subjected to predetermined layout processing such as PDF-formatted document information, at least, (1) Due to processing such as ligatures, it is difficult to handle as simple text data. In addition to the above problems, (2) As illustrated in FIG. 3, when the character data is distributed in two or more regions (A, B) due to stepped arrangements or the like, although the text data should be continuously extracted from the end of region A to region B, the character data may be divided and handled between these regions. (3) When character data is distributed over two or more pages by pagination, the character data may be divided and handled separately. (4) In the process of extracting character data, when attempting to output the extraction result by dividing it at each predetermined delimiter (for example, a period), there are cases where it may be divided at a delimiter part that should not be originally divided (in the case of English text, the dot at the end of an abbreviation not in the dictionary). (5) There are cases where the extraction result is not as intended due to typos etc. originally included in the character data array. (6) If there are parts that do not conform to the rules used in the process of extracting character data, such as parenthetical writing, the character data included in the segment may become long, which may cause problems in subsequent processes such as machine translation. (7) In cases where character data is represented as an image, when extracting the relevant part by OCR (Optical Character Recognition), there are cases where the extraction result is not as intended due to reading errors etc. Problems such as these occur.
[0028] Specifically, the character data extraction unit 22 of the present embodiment divides and outputs character data into segments such as sentence units. However, when the character data continues across the above (2) multiple regions or (3) continues over two or more pages, the division is also performed between regions or between pages. For example, in the first region (FIG. 3(A)), a part on the leading side of the sentence is described, and the part following in the second region (B) transistor structure, a transistor junction 37 between the base layer 36 and the wafer 32 with the upper terminus of this junction lying beneath the oxide mask 31, as illustrated in FIG. 6. is described. When these are described, the character data extraction unit 22 divides them into different segments, pre-translates the original data of each divided segment, and outputs the result.
[0029] That is, in this example, the character data extraction unit 22 … <trans-unit id="1"> <source> In conventional manner there is produced through controlled heating and cooling of the <target>In the conventional method, it is manufactured by controlled heating and cooling < / target> <trans-unit id="2"> <source> transistor structure, a transistor junction 37 between the base layer 36 and the wafer 32 with the upper terminus of this junction lying beneath the oxide mask 31, as illustrated in FIG. 6. <target>Transistor structure, at the transistor junction 37 between the base layer 36 and the wafer 32, as shown in FIG. 6, the upper end of this junction is under the oxide mask 31. < / target> … will output an xliff file such as this. Here <trans-unit id="1">The part surrounded by a segment tag such as corresponds to one segment.
[0030] At this time, the estimation processing unit 23 estimates a series of text data based on the column of character data divided into two or more blocks as a result of the layout processing as described above. That is, the estimation processing unit 23 estimates the division position of the original segment based on the column of character data that is divided into a plurality of segments while including character data that should be continuous as a single segment.
[0031] Specifically, the estimation processing unit 23 sequentially selects each segment included in the xliff file output by the character data extraction unit 22 as a processing target from the beginning, and then performs the following processing. Hereinafter, it is assumed that serial numbers are assigned to each segment in order from the beginning.
[0032] The estimation processing unit 23 selects, as a target range, n + 1 segments from the (p - n)-th to the p-th segments (where n is an integer of 0 or more) that at least include the selected processing target segment (the p-th segment from the beginning), and m segments (where m is an integer of 1 or more) after the p-th segment of the processing target (that is, the segments up to the (p + m)-th segment). Note that n and m may be determined in advance. However, when there are no n preceding segments (that is, when p < n), such as when the processing target segment is the first segment, the segments from the first to the (p + m)-th segments may be selected as the target range. Note that when the number of subsequent segments is less than m, all subsequent segments may be included in the target range and selected.
[0033] The estimation processing unit 23 instructs the generation AI model to concatenate, if possible, consecutive segments among the selected target range segments, or to divide at the parts that should be divided together with the concatenation. This instruction (prompt) is, for example, "In the following English sentences, please correct English sentences that are concatenated before and after without comments and should only output the corrected sentence." It becomes something like this.
[0034] Note that the content in parentheses shows the translation of the instruction when the generative AI model accepts the instruction in English. Also in this example, for the part that says "English sentence" in the prompt, it may vary depending on the language of the sentence represented by the input character data (or the text data to be extracted).
[0035] That is, the estimation processing unit 23 instructs the generative AI model to concatenate the plurality of segments extracted by the character data extraction unit 22 if they are concatenable and to divide them at the positions where they should be divided, based on, for example, the language of the sentence represented by the input character data (or the text data to be extracted), and outputs the segments in the selected attention range.
[0036] Here, the generative AI model may adopt, for example, a large language model or the like that is suitable for the process of determining whether sentences are continuous. By this process, when the selected segment includes the character data of the segment divided between the above regions A and B, that is, when n = 1 and m = 2 and sentence2 is selected as the processing target, the plurality of selected segments are: sentence 1 : Upon the upper surface of the Wafer 32 exposed at … sentence 2 : In conventional manner there is produced through controlled heating and cooling of the sentence 3 : transistor structure, a transistor junction 37 between the base layer 36 and the wafer 32 with the upper terminus of this junction lying beneath the oxide mask 31, as illustrated in FIG. 6. sentence 4 : As silicon technology is … (Sentences 2 and 3 are actually consecutive, but are split because they were extracted from different regions.) The estimation processor 23 concatenates and outputs character data of consecutive segment pairs (sentences 1 and 2, sentences 2 and 3, ...) that form consecutive sentences.
[0037] In the above example, the estimation processing unit 23 connects sentence 2 and sentence 3 through processing using the generative AI model, and obtains sentence 1 : Upon the upper surface of the Wafer 32 exposed at … Sentence 2: In conventional manner there is produced through controlled heating and cooling of the transistor structure, a transistor junction 37 between the base layer 36 and the wafer 32 with the upper terminus of this junction lying beneath the oxide mask 31, as illustrated in FIG. 6. Sentence 4: As silicon technology is … It will be output as …
[0038] The estimation processing unit 23 changes the connected segments into one segment, selects the next segment as the processing target, and continues the processing. In the above example, Sentence 2: In conventional manner … Sentence 4: As silicon technology is … Sentence 5: During or following … Sentence 6: This additional masking 38 … (Since sentence 3 is no longer connected to sentence 2), if some of these multiple segments are connectable, they are connected as the selected segments. The estimation processing unit 23 repeats this processing for each segment.
[0039] Also, in this example, when there is a part in the middle of a segment that should be split by the output of the generated AI model, the estimation processing unit 23 may split one segment into multiple segments.
[0040] Through the above processing, the estimation processing unit 23 in this example can output segments that concatenate sentences that continue across multiple paragraphs or sentences that continue across pages.
[0041] By subjecting the segments, which are the processing results of the estimation processing unit 23, to machine translation processing, the document information processing apparatus 1 in the present embodiment … <trans-unit id="1"> <source> In conventional manner there is produced through controlled heating and cooling of the transistor structure, a transistor junction 37 between the base layer 36 and the wafer 32 with the upper terminus of this junction lying beneath the oxide mask 31, as illustrated in FIG. 6. <target>In the conventional method, by heating and cooling the transistor structure, as shown in FIG. 6, a transistor junction 37 is formed between the base layer 36 and the wafer 32, and the upper end of this junction is under the oxide mask 31.< / target> … will obtain a relatively accurate translation as described above.
[0042] Furthermore, in the extraction process of character data (4), when the character data extraction unit 22 attempts to output each segment for each predetermined delimiter (for example, splitting sentences by periods), it may split at a delimiter part that should not be split originally (in the case of English text, the dot at the end of an abbreviation not in the dictionary).
[0043] Also in this case, through the above processing of the estimation processing unit 23, segments that have been split unintentionally can be concatenated into one segment and output. That is, it becomes possible to estimate and output text data that is a continuous single segment (for example, a sentence) from character data that has been split into multiple segments.
[0044] In addition, when the estimation processing unit 23 uses a generative AI model based on a large language model, even if the extraction result becomes unintended due to typos or disappearance of spaces (delimiters between words) originally included in the arrangement of character data (5), it becomes possible to estimate text data with typos and the like in the character data corrected.
[0045] As in the examples described so far, the estimation processing unit 23 normalizes ligatures and combined character strings (replaces them with characters in a predetermined character set) through processing using, for example, a generative AI model based on a large language model, and also concatenates character data that should originally be in one segment but has been split into two or more segments.
[0046] Furthermore, when there are parts in the character data in a segment that do not conform to the rules used in the extraction process of character data, such as (6) parentheses, and the character data in the segment becomes too long, the estimation processing unit 23 may perform shortening processing to estimate text data that conforms to the gist of the original sentence.
[0047] Specifically, the estimation processing unit 23 in this example selects the character data of each segment (if the output of the character data extraction unit 22 is xliff data, the character string within the source tag of each segment), and uses both the character data of the selected segment and a prompt that instructs to perform shortening of the character data if possible as the input to the generation AI model.
[0048] Specifically, this prompt may be something like "Please split the following English sentence if it contains a complex sentence." In this example, for the part "English sentence" in the prompt, it may vary depending on the language of the sentence represented by the input character data (or the text data to be extracted).
[0049] According to this example, the estimation processing unit 23 can shorten a long sentence through the process using the generation AI model with this prompt input and provide it for subsequent processing. This process is also effective, for example, when an unintended long sentence occurs that the character data extraction unit 22 did not split into segments. This is the case, for example, We will retain your data sent through the API for up to 30 days. [If the data is too large, it may be deleted from the oldest even if it is within 30 days.] On the termination of … For character data containing multiple sentences such as this, even when the character data extraction unit 22 treats it as one segment (when the part in parentheses cannot be extracted and the whole is recognized as one sentence), through the process using the generation AI model, Sentence 1: We will retain your data sent through the API for up to 30 days. Sentence 2: [If the data is too large, it may be deleted from the oldest even if it is within 30 days.] Sentence 3: On the termination of … It can be expected to be divided like this, and when it is actually divided like this, by subjecting each to, for example, machine translation processing, it becomes possible to translate the shortened sentences, and an improvement in translation accuracy can be expected.
[0050] Here, it may be possible to perform this process for all segments. However, the estimation processing unit 23 may perform this process only for segments that satisfy a predetermined condition, such as when the number of characters or words included in the segment exceeds a predetermined threshold.
[0051] Furthermore, in cases where the (7) character data is represented as an image, etc., assuming that the extraction result is unintended due to a reading error, etc., when the character data is extracted by OCR, the estimation processing unit 23 selects the character data of each segment included in the output generated by the character data extraction unit 22 for the OCR result (if the output of the character data extraction unit 22 is xliff data, the character string within the source tag of each segment), and while correcting the typo, etc. included in the selected segment's character data, and if there is a divisible part, generates a prompt instructing to divide it, and uses both as the input to the generation AI model.
[0052] According to this example, the estimation processing unit 23 corrects typos (including reading errors by OCR) through processing using the generation AI model performed by inputting this prompt, and also segments text data obtained by segmenting segments at parts where the character data extraction unit 22 fails to segment, such as cases where there is no space immediately after a period and the next sentence has started. That is, in this example, the estimation processing unit 23 estimates a set of text data re-segmented in sentence units based on the character data of each segment.
[0053] [Operation] The document information processing apparatus 1 of the present embodiment has a configuration as in the above example and operates as in the following example. In the following example, the user inputs, as document information including character data of English text subjected to predetermined layout processing, PDF data including two-tiered character data as illustrated in FIG. 3, to the document information processing apparatus 1 as the translation target. Also, the user will translate the sentences in this PDF data into Japanese.
[0054] Upon receiving the input of the PDF data, the document information processing apparatus 1 starts the processing shown in FIG. 4, extracts the character data included in the input PDF data by a widely known method, and segments it into segments according to a predetermined rule such as sentence units (S11). For this segmentation process, a widely known method in the prior art can be adopted, but in this method, sentences are segmented into different segments at positions straddling the tiers of the PDF data. Here, it is assumed that the direction of the tiers is known, and it is known that the sentences in region B (the right tier) are connected following the sentences in region A (the left tier).
[0055] The document information processing apparatus 1 subjects each segmented segment (including sentences segmented at incorrect positions as described above) to preliminary translation processing and outputs the result (including the original text and the translated text) in the form of xliff (S12). This output result may be displayed and output on the display unit 14.
[0056] Also, the document information processing apparatus 1 executes the following processing while sequentially selecting the segments obtained in step S11 in order from the head. The document information processing apparatus 1 uses, as the input to the generation AI model, the character data within the selected segment (here, the content of the source tag, which is the original text among the data in the xliff format) along with the prompt "In the following English text, concatenate the text excluding the tag parts to correct the English text, and output only the corrected sentence." to execute the estimation process (S13).
[0057] As a result, a character string including ligatures such as signi<0001>fi <!--0001-->cant will be expressed only by characters included in a predetermined character set that does not include ligatures, such as "significant". Further, the output (estimation result) at this time is text data generated by the generation AI model, and even if there are errors such as typos in the character data within the segment, it will be in a corrected state.
[0058] Note that the document information processing apparatus 1 may execute a spell check process or a rule-based character normalization process (for example, elimination of ligatures, elimination of combining characters, etc.) instead of or together with the process of using the generation AI model at this time.
[0059] The document information processing apparatus 1 replaces the content of the source tag within the selected segment with the estimation result in this step S13 (S14).
[0060] Also, the document information processing apparatus 1 selects, as the target range, n + 1 segments from the (p - n)-th to the p-th segments (where n is an integer of 0 or more) including at least the segment selected in step S11 (the p-th segment from the head), and m segments (where m is an integer of 1 or more) after the p-th segment to be processed (that is, the segments up to the (p + m)-th segment).
[0061] Then, the document information processing apparatus 1 inputs the original text data (the content of the source tag) included in the segment selected as the target range in step S15, and inputs a prompt to the generation AI model to instruct it to concatenate the consecutive original text data that can be concatenated among the original text data included in the segment, and obtains its output (S16). This instruction (prompt) is, for example, something like "In the following multiple English sentences, please concatenate those that are connected before and after to correct the English sentence. Comments are not required. Please show only the corrected English sentence."
[0062] By the processing of this step S16, the generation AI model outputs the original text data of one segment by concatenating the original text data that is considered to be consecutive to each other. That is, as a result of the layout processing, the generation AI model estimates the text data in a state where the columns of character data divided into two or more blocks are concatenated as a series of text data.
[0063] The document information processing apparatus 1 compares the text data for each segment included in the output of step S16 and the original text data for each input segment in a forward match in order from the first segment, and in the segment included in the output of step S16 (hereinafter referred to as the output segment for the sake of distinction), replaces the original text data of the segment that matches the text data of the output segment forward (excluding the segment including the original text data determined to match the text data of another output segment that has already preceded) to update the file in xliff format (S17). From this, for example, each segment (denoted as sentence n) included in the target range is sentence 1 : Upon the upper surface of the Wafer 32 exposed at … sentence 2 : In conventional manner there is produced through controlled heating and cooling of the sentence 3 : transistor structure, a transistor junction 37 between the base layer 36 and the wafer 32 with the upper terminus of this junction lying beneath the oxide mask 31, as illustrated in FIG. 6. sentence 4 : As silicon technology is … When sentence 1 : Upon the upper surface of the Wafer 32 exposed at … sentence 2: In conventional manner there is produced through controlled heating and cooling of the transistor structure, a transistor junction 37 between the base layer 36 and the wafer 32 with the upper terminus of this junction lying beneath the oxide mask 31, as illustrated in FIG. 6. sentence 4: As silicon technology is … The xliff file is updated in this way.
[0064] The generative AI model used in the estimation process in step S13 and the generative AI model used in step S16 do not necessarily have to be the same. Furthermore, when character data included in a segment becomes long due to the presence of a part that does not match the rules used in the character data extraction process, such as parentheses, the document information processing device 1 may divide the character data into short sentence segments by processing using the generative AI model already described.
[0065] If there is a segment following the selected segment, the document information processing apparatus 1 selects the segment and repeats the process from step S13. If there is no segment following the selected segment, the document information processing apparatus 1 subjects the XLIFF format file updated in the process from step S13 to step S17 to machine translation processing again. In this machine translation processing, the content of the source tag for each segment included in the updated XLIFF format file is subjected to predetermined processing such as being subjected to machine translation (S18).
[0066] At this time, the content of the source tag for each segment is in a state where ligatures and the like have been resolved, sentences separated by line breaks and pagination have been combined, and typos and the like have also been corrected by the processing from step S13 to step S17, so that the original data is in a state where the accuracy of the machine translation processing can be expected to be improved compared to the preliminary translation processing in step S12.
[0067] Note that the result of the machine translation processing here may be corrected using a generative AI model so that a word in the original language of the specified original text is translated as a specified word in the original language of the translated text.
[0068] In the description so far, the document information processing apparatus 1 segments the character data extracted from the original document information into units such as sentences, and for each of these segments or for a group of these segments, uses a machine learning model (generative AI model) to estimate the original text data, thereby resolving ligatures, resolving typos, and estimating the original sentence of a sentence that has been unintentionally split. However, the present embodiment is not limited to this.
[0069] For example, in the present embodiment, the result of segmentation may be presented to the user, and the user's selection of the segment to be processed may be received, and the processing of steps S13 to S17 described above may be performed on the selected segment.
[0070] According to this embodiment, the manual work involved in correction can be reduced, and the processing efficiency can be improved.
[0071] According to this embodiment, by integrating the generative AI into the translation support tool, it has become possible to perform corrections that were not possible before, such as correcting spelling mistakes in the manuscript according to the context, correcting word concatenation due to missing spaces, correcting typos and omissions, correcting the paragraph layout of pdf documents and the splitting of sentences at page boundaries, correcting segmentation errors due to new abbreviations and segmentation errors where multiple sentences are included in one segment, and furthermore, performing character combination corrections.
[0072] Also, when the manuscript is OCR, particularly obvious typos, omissions, spelling mistakes in the manuscript, misspellings due to missing spaces between words, segmentation errors, etc. will result in mistranslations in machine translation. Therefore, translators needed to perform post-editing work to correct these mistranslations. According to the present invention, the above-mentioned defects in the manuscript are automatically corrected, eliminating the need for translators to perform post-editing work related to the above-mentioned defects, thus enabling faster translation and reducing translation costs.
Description of Reference Numerals
[0073] 1 Document information processing device, 11 Control unit, 12 Storage unit, 13 Operation unit, 14 Display unit, 15 Communication unit, 21 Reception unit, 22 Character data extraction unit, 23 Estimation processing unit, 24 Output unit.
Claims
1. Receiving means for receiving document information including character data subjected to predetermined layout processing; Estimation means for estimating text data to be extracted from the document information based on the character data included in the received document information; Processing means for subjecting the estimated text data to predetermined translation processing; A document information processing apparatus comprising the above.
2. The document information processing apparatus according to claim 1, wherein the estimation means uses a generative AI model which is a large language model based on Generative Pre-trained Transformer to generate a sentence based on an instruction in a language, inputs the character data into the generative AI model, and instructs to estimate the text data to be extracted from the document information, and acquires the estimated text data as an output of the generative AI model.
3. The document information processing apparatus according to claim 2, wherein the estimation means estimates text data in which a ligature or a combined string included in the character data is replaced with a character in a predetermined character set.
4. The document information processing apparatus according to claim 2, wherein the estimation means estimates a segmentation position of segments based on a column of the character data which includes character data to be continuous as a single segment by layout processing and is divided into a plurality of segments.
5. The document information processing apparatus according to claim 2, wherein the estimation means estimates a continuous single sentence from text data divided into a plurality of segments based on the character data.
6. The document information processing apparatus according to any one of claims 2 to 5, wherein the estimation means estimates text data in which a typo included in the character data is corrected.
7. A computer, Receiving means for receiving document information including character data subjected to predetermined layout processing; Estimation means for using a generative AI model in a machine-learned state to generate a sentence based on an instruction in a language, inputting the character data included in the received document information into the generative AI model, instructing to estimate the text data to be extracted from the document information, and estimating the text data to be extracted from the document information as an output of the generative AI model; Processing means for subjecting the estimated text data to a predetermined translation process; A program that functions as.
Citation Information
Patent Citations
Deep-learning based text correction method and device
JP2023071598A