Program, information processing system, and information processing method
The program improves translation efficiency by dividing and comparing original and translated text elements, addressing inefficiencies in existing systems by allowing for direct user feedback and improved accuracy.
Patent Information
- Application Number
- JP2024202745
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-17
- Filing Date
- 2024-11-20
- Publication Date
- 2025-07-30
Smart Images

Figure 2025111375000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a program, an information processing system, and an information processing method.
Background Art
[0002] In Patent Document 1, in a machine translation system, when the output translation does not match the user's intended sentence, the input sentence needs to be frequently corrected or re-entered. Therefore, a plurality of different forward translations obtained by translating a source sentence in a first language into a second language are generated, and for each of the plurality of different forward translations, a plurality of backward translations obtained by translating back into the first language are generated. When a plurality of backward translations are output by an information output device and an operation of selecting one backward translation from the plurality of backward translations is received, the forward translation corresponding to the one backward translation is output.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, there is room for improvement in known technologies such as the prior art. In view of the above circumstances, the present invention aims to provide a more beneficial program and the like.
Means for Solving the Problems
[0005] According to one aspect of the present invention, there is provided a program that causes at least one computer to execute the following steps: In a first reception step, a first file including a first document described in a first language is received. In a first output step, a second file including a first table generated based on the first document and a second table is output. The first table divides the first document into a plurality of first elements and stores each of the plurality of first elements. The second table stores each of a plurality of second elements obtained by translating each of the plurality of first elements into a second language. The first table and the second table are arranged such that the correspondence between each of the plurality of first elements and each of the plurality of second elements is known. In a second reception step, the second file is received. In a second output step, a third file including a second document obtained by extracting a plurality of second elements from the second table included in the second file is output.
[0006] According to one aspect of the present invention, it is possible to provide a more useful information processing system or the like.
Brief Description of the Drawings
[0007]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Mode for Carrying Out the Invention
[0008] [First Embodiment] Hereinafter, embodiments of the present invention will be described. Various characteristic matters shown in the following embodiments can be combined with each other.
[0009] The information processing apparatus 1 that executes the processing of this embodiment includes a communication unit 11, a storage unit 12, a control unit 13, a display unit 14, and an input unit 15. A more specific example of the information processing apparatus 1 is a PC (Personal Computer).
[0010] 2. Information Processing Method In this section, the flow of the information processing method executed by the information processing apparatus 1 will be described. As shown below, the information processing method includes the following steps. The information processing system of the present embodiment includes at least one processor. The processor is configured to execute a program so that each step of the program is performed. The information processing system may be configured by one device, and an example of such a device is the information processing apparatus 1. Note that the order of processing can be appropriately changed, a plurality of processes may be executed simultaneously, or some processes may be omitted.
[0011] 2.1 Outline of the Flow
[0012] The program according to the present embodiment causes at least one computer to execute the following steps. As a first reception step, the control unit 13 receives a first file including a first document described in a first language. As a first output step, the control unit 13 outputs a second file including a first table and a second table. The first table divides the first document into a plurality of first elements and stores each of the plurality of first elements. The second table stores each of the plurality of second elements obtained by translating each of the plurality of first elements into a second language. The first table and the second table are arranged so that the correspondence between each of the plurality of first elements and each of the plurality of second elements can be understood.
[0013] According to such an aspect, when receiving an original file (an example of the first file) describing the original text (an example of the first document), it is possible to output a comparison file (an example of the second file) that compares the elements of the original text (an example of the first element) and the elements of the translated text (an example of the second element). FIG. 1 is a reference diagram. This reference diagram shows an example of the second file.
[0014] 2.2 Details of the Flow
[0015] Hereinafter, an example of the flow of the process executed by the information processing apparatus 1 will be shown.
[0016] As a first reception step, the control unit 13 receives a first file including a first document described in a first language.
[0017] Next, as a presentation step, the control unit 13 presents a translation fee calculated based on the number of characters or words in the first document. According to such an aspect, it is possible to present a translation fee calculated based on the number of characters or words in the original text.
[0018] Next, the control unit 13 divides the first document into a plurality of first elements and stores each of the plurality of first elements in a first table. The first table divides the first document into a plurality of first elements and stores each of the plurality of first elements.
[0019] The width of the first table is equal to or less than half of the width of the display area of the second file, and the width of the second table is determined according to the width of the first table. According to such an aspect, even when the first table and the second table are arranged, it is possible to keep within the width of the display area of the comparison file.
[0020] Next, as a replacement step, before performing translation, the control unit 13 performs a process of replacing at least a part of the description in the first table based on a predetermined first rule. The first rule is a rule for pre-processing of translation. According to such an aspect, since at least a part of the description of the original text is replaced based on the first rule before translation, a translation that better meets the user's requirements can be generated. For example, when the original text is a document related to a patent or a utility model, the first rule corresponding to the description peculiar to the patent or the utility model may be used.
[0021] Next, as an input step, the control unit 13 inputs the first table into the translation system.
[0022] Next, as a second reception step, the control unit 13 receives, from the translation system, a second table obtained by translating the first table into a second language. The second table stores each of a plurality of second elements obtained by translating each of the plurality of first elements into the second language. According to such an aspect, it is possible to input the first table generated based on the original text into the translation system and obtain the translated second table.
[0023] Next, as a replacement step, after performing the translation, the control unit 13 performs a reverse replacement process on at least a part of the descriptions corresponding to the replaced descriptions among the descriptions in the second table.
[0024] Next, as a combining step, the control unit 13 integrates the first table and the second table into one table. Each of the corresponding plurality of first elements and each of the plurality of second elements are arranged side by side on the left and right for each corresponding element. According to such an aspect, it is possible to output a comparison file in which each element of the original text and each element of the translated text are arranged side by side on the left and right for each corresponding element. Since it is easy to compare the elements of the original text and the elements of the translated text, it can satisfy the user.
[0025] Next, as a first output step, the control unit 13 outputs a second file including the first table and the second table. The first table and the second table are arranged so that the correspondence between each of the plurality of first elements and each of the plurality of second elements can be understood.
[0026] Here, as a first output step, the control unit 13 may output a second file in which the first table that has not undergone the replacement process and the second table after the reverse replacement are arranged so as to be comparable for each of the plurality of elements. According to such an aspect, the user can compare the translated text obtained by performing the translation after replacing the original text based on the first rule before the translation with the original text.
[0027] Next, as an execution step, the control unit 13 executes a predetermined process for each description of the plurality of second elements based on each description of the plurality of first elements corresponding to each description of the plurality of second elements and the second rule. The second rule includes a rule for arranging the format of the second document. According to such an aspect, in the comparison file, based on the description of the elements of the original text, a predetermined process can be performed on the elements of the translated text. Even if there are notations or description discrepancies in the elements of the translated text, a predetermined process can be performed based on the description of the original text. For example, when the original text is a document related to a patent or utility model, a predetermined process may be performed in accordance with the rules of the country where the translated text is to be submitted after translation.
[0028] Next, as an insertion step, the control unit 13 inserts a claim number or paragraph number in a predetermined format for each element of the second table. If there is no claim number or paragraph number in a predetermined format, a claim number or paragraph number in a predetermined format is inserted for each element. According to such an aspect, when the original text is a document related to a patent or utility model, a claim number or paragraph number in a predetermined format can be inserted into the translated text.
[0029] Next, as a second output step, the control unit 13 outputs a third file including a second document obtained by extracting a plurality of second elements from the second table. According to such an aspect, when receiving the original text file, the translated text (second document) can be output from the generated comparison file. When the control unit 13 receives an output instruction for the translated text, it may be configured to output the translated text.
[0030] The first document is a document related to a patent or utility model. Each of the plurality of first elements is obtained by dividing the first document into claim or paragraph units. According to such an aspect, when the original text is a document related to a patent or utility model, the original text and the translated text can be compared for each claim or each paragraph. The ability to compare the original text and the translated text for each claim or each paragraph is convenient, for example, when considering procedures for foreign countries.
[0031] The above-described mode of information processing is merely an example. The present invention is not limited thereto and can be appropriately modified without departing from the technical idea of the invention.
[0032] (1) A program that causes at least one computer to execute the following steps. In the first reception step, a first file including a first document described in a first language is received. In the first output step, a second file including a first table and a second table is output. The first table divides the first document into a plurality of first elements and stores each of the plurality of first elements. The second table stores each of the plurality of second elements obtained by translating each of the plurality of first elements into a second language. The first table and the second table are arranged such that the correspondence between each of the plurality of first elements and each of the plurality of second elements is understandable.
[0033] (2) In the program according to (1) above, further, in the combining step, the first table and the second table are integrated into one table. Here, each of the plurality of corresponding first elements and each of the plurality of second elements are arranged side by side on the left and right for each corresponding element.
[0034] (3) In the program according to (1) above, further, in the input step, the first table is input to a translation system. Further, in the second reception step, the second table obtained by translating the first table into the second language is received from the translation system.
[0035] (4) In the program according to (3) above, the width of the first table is equal to or less than half of the width of the display area of the second file, and the width of the second table is determined according to the width of the first table.
[0036] (5) In the program according to (1) above, further, in the replacement step, before translation, a process of replacing at least a part of the description in the first table is performed based on a predetermined first rule, and the first rule is a rule for preprocessing translation. Program.
[0037] (6) In the program according to (5) above, in the replacement step, after translation, a reverse replacement process is performed on at least a part of the description corresponding to the replaced description among the descriptions in the second table. In the first output step, a second file in which the first table not subjected to the replacement process and the second table after reverse replacement are arranged so as to be comparable for each of a plurality of elements is output. Program.
[0038] (7) In the program according to (1) above, further, in the second output step, a third file including a second document obtained by extracting the plurality of second elements from the second table is output. Program.
[0039] (8) In the program according to (7) above, further, in the execution step, based on the description of each of the plurality of first elements corresponding to the description of each of the plurality of second elements and a second rule, a predetermined process is performed on the description of each of the plurality of second elements. The second rule includes a rule for arranging the format of the second document. Program.
[0040] (9) In the program according to (1) above, the first document is a document related to a patent or a utility model, and each of the plurality of first elements is obtained by dividing the first document into claim or paragraph units. Program.
[0041] (10) In the program according to (9) above, further, in the insertion step, if each of the elements in the second table does not have a claim number or a paragraph number in a predetermined format, the claim number or the paragraph number in the predetermined format is inserted into each of the elements. Program.
[0042] (11) In the program according to (1) above, further, in the presentation step, a program that presents a translation fee calculated based on the number of characters or words in the first document.
[0043] (12) An information processing system comprising at least one processor, wherein the processor is configured to execute the program so that each step of the program according to any one of (1) to (11) above is performed.
[0044] (13) An information processing method comprising each step of the program according to any one of (1) to (11) above.
[0045] [Second Embodiment] Hereinafter, a second embodiment of the present disclosure will be described with reference to the drawings. Various characteristic matters shown in the embodiments described below can be combined with each other. Also, the first embodiment described above and the second embodiment described below can be combined as appropriate.
[0046] By the way, a program for realizing software appearing in an embodiment may be provided as a non-transitory computer-readable medium readable by a computer, may be provided so as to be downloadable from an external server, or may be provided so that the program is started on an external computer and its function is realized on a client terminal (so-called cloud computing).
[0047] Also, in various information processes according to an embodiment, an input and an output corresponding to the input can be realized. Here, if an output is obtained as a result of the input, the form of information (hereinafter referred to as reference information) referred to in such information processing is not limited. The reference information may be, for example, rule-based information such as a database, a lookup table, a predetermined function (including determination expressions such as regression equations constructed by statistical methods), a learned model in which the correlation between the input and the output has been learned in advance, or a large language model capable of outputting a desired result by inputting a prompt.
[0048] Also, in one embodiment, the "unit" may include, for example, a combination of hardware resources implemented by a circuit in a broad sense and information processing of software that can be specifically realized by these hardware resources. Also, in one embodiment, various information is handled, and these information are represented, for example, by physical values of signal values representing voltage and current, the high and low of signal values as a set of binary bits composed of 0 or 1, or quantum superposition (so-called quantum bits), and communication and calculation can be executed on a circuit in a broad sense.
[0049] Furthermore, a circuit in a broad sense is a circuit realized by appropriately combining at least a circuit, circuitry, a processor, and a memory, etc. Also, the processor may be a general-purpose processor or a dedicated circuit. That is, it includes application specific integrated circuits (ASICs), programmable logic devices (for example, simple programmable logic devices (SPLDs), complex programmable logic devices (CPLDs), and field programmable gate arrays (FPGAs)), etc.
[0050] 1. Hardware Configuration In this section, the hardware configuration of the information processing system according to this embodiment will be described.
[0051] <Information Processing System 100>
[0052] The information processing system 100 constitutes a part of a translation support system that can comprehensively support each process of the translation business, such as the translation of the original text, the creation of the final delivered product, and the calculation of translation fees. In one embodiment, the information processing system 100 is composed of one or more devices or components. FIG. 2 is a diagram showing the overall configuration of the information processing system 100. The information processing system 100 includes a server 10, a translation system 20, and one or more user terminals 30. The server 10, the translation system 20, and the user terminals 30 are configured to be communicable through a network. The connections of the server 10, the translation system 20, and the user terminals 30 to the network may be wired or wireless.
[0053] <Server 10> Next, the server 10 will be described with reference to FIG. 3. FIG. 3 is a block diagram showing the hardware configurations of the server 10 and the translation system 20 (FIGS. 3(a) and 3(b)). The server 10 is an information processing device that executes processes for supporting the translation business, such as formatting the first document, which is the original text to be translated, and the document obtained by automatically translating the first document (hereinafter also simply referred to as the "primary translation text"), while communicating with the translation system 20 and the user terminals 30 via the network. As shown in FIG. 3(a), the server 10 includes a control unit 101, a storage unit 102, a communication unit 103, and a communication bus 104. The control unit 101, the storage unit 102, and the communication unit 103 are electrically connected inside the server 10 via the communication bus 104.
[0054] <Control Unit 101> The control unit 101 processes and controls the overall operations related to the server 10. The control unit 101 is, for example, a Central Processing Unit (CPU). The control unit 101 realizes various functions related to the server 10 by reading a predetermined program stored in the storage unit 102. That is, the information processing by software stored in the storage unit 102 is specifically realized by the control unit 101, which is an example of hardware, and can be executed as each functional unit in the control unit 101. These will be described in more detail in the next section. Note that the control unit 101 is not limited to being single, and the server 10 may have a plurality of control units 101 for each function. Also, the server 10 may be configured by a combination of these.
[0055] <Storage unit 102> The storage unit 102 stores various information defined by the foregoing description. This can be implemented, for example, as a storage device such as a Solid State Drive (SSD) that stores various programs and the like related to the server 10 executed by the control unit 101, or as a memory such as a Random Access Memory (RAM) that stores temporarily necessary information (arguments, arrays, etc.) related to the calculation of programs. The storage unit 102 stores various programs, variables, etc. related to the server 10 executed by the control unit 101.
[0056] <Communication unit 103> The communication unit 103 may be a wired communication means such as USB, IEEE1394, Thunderbolt (registered trademark), wired LAN network communication, etc., or may be a wired communication means such as mobile communication of 3G / LTE / 5G, BLUETOOTH (registered trademark) communication, wireless LAN network communication, etc. Also, the communication unit 103 is preferably implemented as a set of these plural communication means. That is, the server 10 may communicate various information with the outside via the communication unit 103 and the network.
[0057] The server 10 may be in an on-premises form or in a cloud form. The server 10 in the cloud form may provide the above-described functions and processes in forms such as, for example, SaaS (Software as a Service) and cloud computing.
[0058] <Translation system 20> The translation system 20 is an information processing device that executes translation processing, such as generating a primary translation by translating a first document received from the server 10 via a network and transmitting the file including the primary translation to the server 10 via the network. As shown in FIG. 3(b), the translation system 20 includes a control unit 201, a storage unit 202, a communication unit 203, and a communication bus 204. The control unit 201, the storage unit 202, and the communication unit 203 are electrically connected via the communication bus 204 inside the translation system 20. The descriptions of the control unit 201, the storage unit 202, and the communication unit 203 are omitted because they are the same as the descriptions of the respective parts in the server 10.
[0059] <User terminal 30> FIG. 4 is a block diagram showing the hardware configuration of the user terminal 30. As shown in FIG. 4, the user terminal 30 includes a control unit 301, a storage unit 302, a communication unit 303, a display unit 304, an input unit 305, and a communication bus 306. The control unit 301, the storage unit 302, the communication unit 303, the display unit 304, and the input unit 305 are electrically connected via the communication bus 306. The descriptions of the control unit 301, the storage unit 302, and the communication unit 303 are omitted because they are the same as the descriptions of the respective parts in the server 10.
[0060] <Display unit 3********** The display unit 304 displays a screen of a graphical user interface (GUI) operable by the user. The display unit 304 may be included in the housing of the user terminal 30 or may be externally attached. Specifically, the display unit 304 can be implemented as a display device such as a CRT display, a liquid crystal display, an organic EL display, or a plasma display. It is preferable that these display devices are selectively implemented according to the type of the user terminal 30.
[0061] <Input unit 305> The input unit 305 receives operation inputs made by the user. The operation inputs are transferred as command signals to the control unit 301 via the communication bus 306. The control unit 301 can execute predetermined control and calculations based on the transferred command signals as necessary. The control unit 301 may be included in the housing of the user terminal 30 or may be externally attached. For example, the input unit 305 may be integrated with the display unit 304 and implemented as a touch panel. When the input unit 305 is implemented as a touch panel, the user can input a tap operation, a swipe operation, etc. to the input unit 305. As the input unit 305, instead of a touch panel, a switch button, a mouse, a track pad, a QWERTY keyboard, etc. can be adopted.
[0062] 2. Functional configuration In this section, the functional configuration of this embodiment will be described. The information processing by software stored in the storage unit 102 is specifically realized by the control unit 101 which is an example of hardware, and thus can be executed as each functional unit included in the control unit 101 (at least one processor included in the information processing system 100). FIG. 5 is a block diagram showing the functions realized by the server 10 (control unit 101).
[0063] As shown in FIG. 5, the server 10 (control unit 101) includes a reception unit 111, a generation unit 112, a replacement unit 113, an inspection unit 114, an insertion unit 115, an integration unit 116, an execution unit 117, an extraction unit 118, a calculation unit 119, and an output unit 110.
[0064] <Reception unit 111> The reception unit 111 is configured to receive various data from the storage unit 102, or from the translation system 20, the user terminal 30, or other information processing terminals as a reception step. For example, the reception unit 111 is configured to receive information such as a first file including a first document, a file including a primary translation, rules related to information processing, and various requests from the translation system 20 or the user terminal 30.
[0065] <Generation unit 112> The generation unit 112 is configured to generate various data, particularly tables and files, as a generation step. For example, the generation unit 112 is configured to generate a first table based on the first document.
[0066] <Replacement unit 113> The replacement unit 113 is configured to replace a predetermined description included in a document with another description as a replacement step. For example, the replacement unit 113 is configured to replace the description included in the first document based on the first rule described later.
[0067] <Inspection unit 114> The inspection unit 114 is configured to inspect whether there is a description in a predetermined format in a document as an inspection step. For example, the inspection unit 114 is configured to inspect whether there is a number in a predetermined format in a second table obtained by translating the first table.
[0068] <Insertion unit 115> The insertion unit 115 is configured to insert a description in a predetermined format into a document as an insertion step. For example, when the inspection result by the inspection unit 114 indicates "not present", the insertion unit 115 is configured to insert a number in a predetermined format into the description of the second table.
[0069] <Integration unit 116> The integration unit 116 is configured to combine a plurality of tables as a combining step. For example, the integration unit 116 is configured to combine the first table and the second table.
[0070] <Execution unit 117> The execution unit 117 is configured to arrange the format of the document as an execution step. For example, the execution unit 117 is configured to arrange the description included in the second table based on the third rule described later.
[0071] <Extraction unit 118> The extraction unit 118 is configured to extract the description in the document as an extraction step. For example, the extraction unit 118 is configured to extract the description included in the second table.
[0072] <Calculation unit 119> The calculation unit 119 is configured to perform various calculations as a calculation step. For example, the calculation unit 119 is configured to calculate the translation fee based on the number of words or characters in the first document.
[0073] <Output unit 110> The output unit 110 is configured to output various data to the storage unit 102, or to the translation system 20, the user terminal 30, or other information processing terminals as an output step. For example, the output unit 110 is configured to output a file including the first table and a file including various documents with the format arranged to the translation system 20 or the user terminal 30.
[0074] 3. Information processing flow In this section, the flow of information processing (information processing method) executed by the information processing system 100 will be described. In this information processing, the program stored in the storage unit 102 is read by the control unit 101 which is a processor, and the following steps are executed. Note that the order of processing can be changed as appropriate, multiple processes can be executed simultaneously, or some processes can be omitted.
[0075] 3.1 Overview First, with reference to FIG. 6, the overview of the flow of information processing according to this embodiment will be described. FIG. 6 is a flowchart showing the overview of the information processing executed by the information processing system 100.
[0076] First, as a first reception step, the reception unit 111 receives a first file including a first document described in a first language (step S101). Subsequently, as a first output step, the output unit 110 outputs a second file including a first table and a second table generated based on the first document (step S102). Note that the first table divides the first document into a plurality of first elements and stores each of the plurality of first elements. The second table stores each of the plurality of second elements obtained by translating each of the plurality of first elements into a second language. The first table and the second table are arranged so that the correspondence between each of the plurality of first elements and each of the plurality of second elements can be understood. Next, as a second reception step, the reception unit 111 receives the second file (step S103). Further, as a second output step, the reception unit 111 outputs a third file including a second document obtained by extracting a plurality of second elements from the second table included in the second file (step S104).
[0077] According to such an aspect, when receiving the original text file (the first file) that describes the original text (the first document), it is possible to output a comparison file (the second file) that compares the elements of the original text (the first element) and the elements of the translated text (the second element). Then, when receiving the comparison file (the second file), it is possible to output a translated text file (the third file) that includes the translated text (the second document) from which only the elements of the translated text (the second element) are extracted.
[0078] That is, the information processing system 100 outputs a third file that includes a second document which is a translated text, based on the first file that includes the first document. According to the information processing system 100, various operations related to the translation work by the user can be facilitated.
[0079] 3.2 Details Hereinafter, with reference to FIGS. 7 to 9, a detailed specific example of the information processing flow according to the present embodiment will be described. FIG. 7 is an activity diagram showing the information processing flow executed by the information processing system 100. FIG. 8 is an activity diagram showing the information processing flow executed by the information processing system 100. FIG. 9 is an activity diagram showing the information processing flow executed by the information processing system 100.
[0080] The specific example may be included within the scope defined by the above-described outline. In the information processing method shown in FIGS. 6 to 9, at least one processor included in the information processing system 100, such as the control unit 101, the control unit 201, or the control unit 301, reads a program stored in the storage unit 102, the storage unit 202, or the storage unit 302, whereby each step of the information processing method is executed. That is, the information processing method includes each step of the program. Further, the program causes at least one computer to execute each step of the information processing method. The activities of the following information processing method may be omitted, repeated, or added according to the embodiment, and some activities may be executed in a different order or a plurality of activities may be executed simultaneously.
[0081] In the following description, as an example, the following situation is assumed. First, an administrator who manages server 10 provides a service capable of generating a third file to users via the Internet. The website that provides this service is referred to as a specific site. That is, this service is a type of SaaS using the specific site, and the administrator is the provider of this service. Various end-users using user terminal 30 are using this service. In this service, when a user accesses the specific site and inputs a first file into server 10, ultimately, a third file, which is a translated file of the first file, is output. The user can download the generated third file and a second file generated as a previous stage of the third file to their own user terminal 30.
[0082] Also, as an example, the case where the first document is a document related to a patent or utility model will be described. Thus, when the first document is a document related to procedures such as a document related to a patent or utility model, by using the information processing system 100 of the present embodiment, it is convenient, for example, when considering responses to procedures in foreign countries. Also, as an example, the case where the first language is English and the second language is Japanese will be described. Note that the first document does not necessarily have to be a document related to a patent or utility model, and the information processing system 100 can be used for translation of various documents. Also, the first language and the second language are not particularly limited to English and Japanese, and any different languages may be used.
[0083] First, the user selects a first file including a first document via the input unit 305 of the user terminal 30 and inputs an instruction to generate a second file (Activity A101). As described above, the first document is the original text to be translated described in the first language, and in this embodiment, it is a document related to a patent or utility model described in English. The first file and the generation instruction are transmitted to the server 10 via the network. Subsequently, the reception unit 111 of the server 10 receives, as a first reception step, the first file and the generation instruction for the second file (Activity A102).
[0084] Subsequently, the generation unit 112 of the server 10, as a first generation step, divides the first document included in the first file into a plurality of first elements, and generates a first table by storing each of the plurality of first elements. That is, the generation unit 112 generates a first table based on the first document (Activity A103).
[0085] Here, each of the plurality of first elements is obtained by dividing the first document into units such as, for example, one chapter, one page, one item, one paragraph, or one sentence. In this embodiment, it is divided into units of claims or paragraphs. The generation unit 112 may divide the first document into a plurality of first elements by recognizing these units based on reference information specifying, for example, format, page break, line break, leading space, period, or full stop, or by inputting the first file to reference information such as a machine learning model to divide the first document into a plurality of first elements.
[0086] Next, the replacement unit 113, as a first replacement step, executes a process of replacing (or converting) at least part of the descriptions in the first table based on a predetermined first rule before translation (Activity A104). The first rule is a rule for pre-processing of translation and is stored in the storage unit 102, the storage unit 202, or the storage unit 302. The first rule will be described in detail later.
[0087] Subsequently, as a request step, the output unit 110 requests the user terminal 30 to specify a second rule (activity A105). That is, the output unit 110 outputs information requesting the specification of the second rule. Such information is transmitted to the user terminal 30 via the network, and a screen indicating such information is displayed on the display unit 304 (activity A106). Thereby, the user can be prompted to specify the second rule. Note that the second rule is a rule for replacing (or converting) a specified term (first description) described in English (first language) in the first table with a description including a translated term (second description).
[0088] For example, the specified term is a word, term, phrase, sentence, or a combination thereof described in English (first language). The translated term is, for example, a description obtained by converting the specified term into Japanese (second language). Also, for example, the description including the translated term (second description) preferably includes both the specified term (first description) and the translated term with an identifier attached thereto (a description obtained by attaching an identifier to the description obtained by converting the first description into the second language). The identifier is, for example, a symbol, and it is preferable to adopt a symbol that is difficult to be included in the first document. The second rule will be described in more detail later.
[0089] The user can input the specification or cancellation of the second rule via the input unit 305. When such an input is made, the reception unit 111 of the server 10 receives the specification or cancellation of the second rule from the user terminal 30 as a second rule reception step (activity A107). When the second rule is specified, the process proceeds to activity A108, and as part of the first replacement step, the replacement unit 113 executes a process of replacing (or converting) at least a part of the description in the first table based on the specified (predetermined) second rule before performing the translation. On the other hand, when the specification of the second rule is cancelled, the control unit 101 skips activity A108.
[0090] Subsequently, in activity A109 of FIG. 8, the output unit 110 outputs (inputs to the translation system 20) a first table (the first table generated based on the first document) to the translation system 20 as an input step. When the first table is output, it is transmitted to the translation system 20 via the network.
[0091] Subsequently, the control unit 201 of the translation system 20 executes a translation process on the first table as a translation step to generate a second table including a primary translation of the first document in the first table into Japanese (the second language) (activity A110). Here, it is preferable that the control unit 201 executes the translation process while maintaining the format of the first table, whereby a second table storing each of a plurality of second elements can be obtained without requiring complicated processing. Each of the plurality of second elements is a translation of each of the plurality of first elements into the second language as described above.
[0092] When the second table is generated, the reception unit 111 of the server 10 receives the second table from the translation system 20 as a translation text reception step (activity A111). Subsequently, the inspection unit 114 inspects, as an inspection step, whether a claim number or a paragraph number in a predetermined format exists in each of the elements of the second table received by the reception unit 111 at a location where it is required (activity A112). For example, in the case of the present embodiment, it is inspected whether a corner bracket and "Claim ○" (○ is an integer) within the corner bracket exist at the beginning of each claim described in the "Claims" item. Also, it is inspected whether a corner bracket and "0△×□" (△, ×, and □ are integers) within the corner bracket exist at the beginning of each paragraph described in the "Specification" item. If a claim number or a paragraph number in a predetermined format does not exist at a required location, the process proceeds to activity A113.
[0093] In Activity A113, as an insertion step, in the second table, for each element at a location where it is determined that there is no claim number or paragraph number, a claim number or paragraph number in a predetermined format is inserted. In this way, if necessary, by determining the presence or absence of numbers in a predetermined format and performing the insertion process according to the determination result, it is possible to shorten the time when the user edits the translated text in the activities described later. In particular, it is advantageous when there are rules for the number format, such as when the first document to be translated is a document related to a patent or utility model as in this embodiment. On the other hand, when a claim number or paragraph number in a predetermined format exists in a necessary location, Activity A113 is skipped.
[0094] Subsequently, as a second replacement step, the replacement unit 113 performs a reverse replacement process on at least a part of the description corresponding to the description replaced (or converted) based on the first rule or the second rule among the descriptions in the first table after translation (Activity A114). At this time, it is preferable that the replacement unit 113 restores the first table to the first table on which the replacement process has not been performed. This reverse replacement process may be executed, for example, by referring to the first rule or the second rule again.
[0095] Next, as a combining step, the integration unit 116 integrates the first table and the second table into one table (Activity A115). Here, the first table and the second table are arranged so that the correspondence between each of the plurality of first elements and each of the plurality of second elements can be understood. Specifically, it is preferable that each of the plurality of corresponding first elements and each of the plurality of second elements are arranged side by side on the left and right for each corresponding element.
[0096] That is, it is preferable that a paragraph (second element) translated from that paragraph (first element) into Japanese (second language) be placed to the left or right of the paragraph (first element) written in English (first language). Thereby, it is possible to output a second file in which each element of the first document that is the original text and each element of the second document that is the translated text are arranged side by side, left and right, for each corresponding element. This makes it easier to compare the elements of the original text and the translated text, which is convenient when the user later edits the second file. As described above, the unit of the element is not limited to a paragraph.
[0097] Here, the first table to be integrated includes the first document restored by inverse substitution or the like in activity A114. That is, it is the first table before the substitution process is performed (the state before the substitution process). That is, in the integrated single table, the first table including the original first document received in activity A102 and the second table are arranged so as to be comparable for a plurality of elements. Thereby, the user can compare the first document initially selected in activity A101 with the translated text translated after performing the substitution process on the original text based on the first rule or the second rule.
[0098] Then, as a first output step, the output unit 110 outputs a second file including the single table integrated by the integration unit 116 to the user terminal 30 (activity A116). Such a second file is transmitted to the user terminal 30 via a network, and a screen showing the second file is displayed on the display unit 304 (activity A117). Thereby, the user can visually recognize the second file. The user who has visually recognized the second file operates the input unit 305 to edit the second file as necessary. At this time, in the second file, the first table including the first document and the second table including the primary translated text are arranged so that the corresponding relationship is understandable for each corresponding element. Therefore, the user can immediately grasp which part of the primary translated text should be edited (corrected) while referring to the first document.
[0099] When the editing work on the second file is completed or has reached a certain stage, the user can input a formatting instruction for the second file via the input unit 305. When such an input is made, the receiving unit 311 of the server 10 receives the second file and the formatting instruction for the second file as a third receiving step (activity A118). Next, as an execution step, the execution unit 117 executes a formatting process (predetermined process) for each description of a plurality of second elements, which are translations included in the second file (activity A119). This process is executed based on each description of a plurality of first elements corresponding to each description of a plurality of second elements and a third rule. The third rule includes a rule for formatting the second table.
[0100] Examples of the third rule include a rule that requires the description of the patent claims or utility model claims to be repeated at the end of the specification, or a rule that requires the written content of claim 1 to be used as the solution in the abstract. In this case, processing is performed based on each description of the first element. Therefore, even if the description of the second element contains variations in notation due to the translation step or user editing, the specified formatting processing can be performed. For example, as long as the first document contains a description that serves as a guide for this processing, the second element can be formatted even if the description of the corresponding second element varies depending on the processing or user. In particular, when the first document is a document related to a patent or utility model, as in this embodiment, and the original and translated texts follow certain rules, this processing is likely to be performed without any problems. Furthermore, performing this processing can significantly reduce the effort required for translation.
[0101] Next, the output unit 110 outputs the formatted second file to the user terminal 30 (activity A120). The second file is transmitted to the user terminal 30 via the network, and a screen showing the second file is displayed on the display unit 304 (activity A121).
[0102] As a result, the user can further edit the second file or perform a final inspection as needed. When the editing work on the second file is completed or paused, the user can instruct the formatting process for the second file again. In this case, the process returns to activity A118.
[0103] On the other hand, if it is determined that no further editing work on the second file is required, the user can input an instruction to output the third file. When such an input is made, the reception unit 111 of the server 10 receives the second file and the instruction to output the third file from the user terminal 30 (activity A122). Subsequently, as an extraction step, the extraction unit 118 generates a second document, which is a translation, by extracting a plurality of second elements from the second table (activity A123). That is, the extraction unit 118 extracts only the elements (second elements) of the translation from the second table including the elements of the original text (first elements) and the elements of its translation (second elements) to generate a translation.
[0104] Subsequently, the generation unit 112 generates the third file by executing a format change process (predetermined process) on the second document (description of each of the plurality of second elements) based on the fourth rule (activity A124). The fourth rule includes a rule for changing the format of a predetermined description in the second document (plurality of second elements). Examples of the predetermined description include item names and paragraph numbers. In the fourth rule, for these descriptions, the format such as the boldness, color, type, and size of the font is changed. This visually arranges the second document including the second elements, which is a translation, to be easier to read.
[0105] Furthermore, as a fee calculation step, the calculation unit 119 calculates a translation fee based on the number of characters or words in the first document (Activity A125). For example, the calculation unit 119 obtains the number of words in the first document based on the first document included in the second file received in Activity A135, and multiplies the number of words by a predetermined coefficient to calculate the translation fee.
[0106] In addition, the calculation unit 119 calculates the translation fee based on the number of drawings or tables. For example, the calculation unit 119 obtains the number of drawings or tables based on the number of types of text information such as "FIG. ♪" or "Table ☆" (♪ and ☆ are integers or an integer and an alphabet) in the first document, and multiplies the number of drawings or tables by a predetermined coefficient. Then, the final translation fee is calculated by adding the translation fee based on the number of words and the translation fee based on the number of drawings or tables. When the first language is Japanese, that is, when the first document is written in Japanese, the number of characters may be used instead of the number of words, and the number of drawings or tables may be obtained based on the description of "Figure ♪" or "Table ☆" (♪ and ☆ are integers or an integer and an alphabet) enclosed in parentheses.
[0107] Subsequently, as a second output step, the output unit 110 outputs a third file including the second document, which is the translated text, to the user terminal 30, and as a presentation step, outputs (presents) a file (or information) indicating the translation fee (Activity A126). Of course, these files may be output separately, but it is more preferable to output them simultaneously. As a result, on the user terminal 30, it is possible to set so that the third file and the file indicating the translation fee are saved in the same hierarchy at the same time. Therefore, for the user, the correspondence between the third file and the file indicating the translation fee is easy to understand.
[0108] As described above, in the information processing system 100 according to the present embodiment, the processor reads a program to execute an information processing method. In the information processing method, by inputting a first file including a first document as the original text to the server 10, a second file including a first table and a second table is output. The first table and the second table store the correspondence between each of the first elements as the original text and each of the second elements as the primary translation text so that it can be understood. Subsequently, by inputting the second file edited as necessary to the server 10 together with a formatting instruction, a second file with formatted layout is output. Further, by inputting the second file edited as necessary to the server 10 together with an instruction to output a third file, a third file including a second document and a translation fee are output. The second document is obtained by extracting each of the second elements that are the translation texts included in the second file and arranging the format.
[0109] In this way, in the present embodiment, by outputting the second file, the correspondence between the original text (first element) and the primary translation text (second element) can be made easier to view. Also, before generating the primary translation text that is segmented into the second elements (translating the first document), various replacement processes are executed, so the quality of the second elements included in the second file is improved. As a result, the speed of operations such as translation editing or confirmation is improved, and the burden on the user when performing translation work can be reduced. Further, the formatting of the second file and the extraction of the second elements from the second file can be executed. Therefore, it is not necessary for the user to manually create the third file including the second document that is the translation text for final delivery. That is, the translation text can be provided without consuming the limited time and labor of the user in simple operations. Further, since the translation fee can also be presented together with the third file, the user can obtain the billing amount easily and without fear of human error.
[0110] Also, in this embodiment, the server 10 outputs (presents) the file processed by it to the user terminal 30 each time, and returns it. As a result, the user can perform an editing operation on the user terminal 30. Thereby, the user can comfortably perform an editing operation using a familiar word processing software or the like. Also, when the server 10 is provided separately from the user terminal 30 as in this embodiment, even if the network connection status is poor, there is no hindrance to the editing operation. Further, even if a problem occurs in the server 10, there is no problem with the editing operation itself.
[0111] 4. Details of Information Processing In this section, the detailed part of the information processing outlined in the previous section will be described with the use of separate figures and the like.
[0112] <First File 4> FIG. 10 is a diagram showing an example of the first file 4. The first file 4 includes a first document 41. In this embodiment, the first document 41 includes text information regarding a patent or utility model described in English (the first language). In the activity A101 of FIG. 7, when the output unit 314 of the user terminal 30 outputs (inputs to the server 10) the first file 4 including this first document 41 and an instruction to generate a second file to the server 10, in the activity A102, the reception unit 111 of the server 10 receives these. Examples of the instruction to generate a second file include an input operation on a UI component such as a button (not shown).
[0113] <First Table 5> In the activity A103 of FIG. 7, the generation unit 112 of the server 10 generates a first table as shown in FIG. 11 based on the first document 41 of FIG. 10 received by the reception unit 111. FIG. 11 is a diagram showing an example of the first table 5. The text information included in the first document 41 is divided by claim or paragraph units, thereby generating a plurality of first elements 52. Then, these plurality of first elements 52 (first elements 52a to 52i) are stored in a plurality of columns 51 (columns 51a to 51i) of the first table 5, respectively.
[0114] <Processing Based on the First Rule> In activity A104 of FIG. 7, when substitution unit 113 executes substitution processing based on the first rule as the first substitution step, the first table becomes as shown in FIG. 12. FIG. 12 is a diagram showing an example of the first table 6 after executing the first substitution step. Among the first elements 52a to 52i stored in columns 51a to 51i of the first table 5 in FIG. 11, the first elements 52g and 52i stored in columns 51g and 51i are stored as first elements 62h and 62j in columns 61h and 61j of the first table 6 in FIG. 12 as they are.
[0115] The first elements 52a, 52e, 52f, 52h stored in columns 51a, 51e, 51f, 51h of FIG. 11 include text information indicating item names in English (the first language), such as "Claims", "What is claimed is:", "Title of invention", "FIELD", and "BACKGROUND". On the other hand, in the first elements 62a, 62f, 62g, 62i stored in columns 61a, 61f, 61g, 61i of FIG. 12, these text information are replaced (or converted) with text information in Japanese (the second language), such as "[Document Name] Claims", "[Name of Invention]", "[Technical Field]", and "[Background Art]" (the angle brackets are corner brackets in FIG. 12).
[0116] Also, the first elements 52b to 52d stored in columns 51b to 51d of FIG. 11 include text information consisting of combinations of numbers and periods, such as "1.", "2.", and "3.", at the beginning. In the first elements 62b to 62d stored in columns 61b to 61d of FIG. 12, these text information are replaced with text information consisting of Japanese (the second language), numbers, and corner brackets, such as "[Claim 1]", "[Claim 2]", and "[Claim 3]" (the angle brackets are corner brackets in FIG. 12).
[0117] Furthermore, in the first table 6 of FIG. 12, a column 61e that did not have a corresponding one in the first table 5 of FIG. 11 is inserted, and a first element 62e, which is text information of "[Document Name] Specification" (the angle brackets are corner brackets in FIG. 12), is stored in this column 61e. The first rule is thus a rule for performing a replacement or conversion process on the first table.
[0118] In this way, at least a part of the text information described in the first language (English and numbers) in the first document, which is the original text, is replaced with text information described in the second language (Japanese and numbers) in advance before translation. Also, if necessary, a column including a first element that does not exist in the first document as the original text is added between the columns of the first table, or a column including a specific first element is deleted. Thereby, a translated text that better meets the user's requirements can be generated. This is particularly useful when the second document needs to include predetermined descriptions, for example, when the first document is a document related to a patent or a utility model.
[0119] The replacement unit 113 may change the first rule based on the translated second language. Further, in an environment where the information processing system 100 is used, when only translation between two specific languages (for example, Japanese and English) is performed, it is preferable to change the first rule based on the first language of the original text. That is, for example, when only Japanese-English translation is performed and the first language is English, the first rule set on the assumption that the second language is Japanese is used, and when the first language is Japanese, the first rule set on the assumption that the second language is English may be used. Also, the first rule may be selected based on the type or content of the first document. For example, when the first document is a document related to a patent or utility model and the second language is Japanese, the first rule for replacing a part of the first element with the description defined by the Japan Patent Office is used, and when the second language is English, the first rule for replacing a part of the first element with the description defined by the United States Patent and Trademark Office may be used. By configuring in this way, for example, in an environment where the first document of various contents or fields is to be translated, a translated file (third file) in the required format can be output.
[0120] <File 7 including the second rule> In response to the output unit 110 of the server 10 requesting the user terminal 30 to specify the second rule (Activity A105 in FIG. 7), a file including the second rule as shown in FIG. 13 is specified (Activity A106), and such a file is received by the reception unit 311 (Activity A107). FIG. 13 is a diagram showing an example of File 7 including the second rule (FIG. 13(a)) and a diagram showing an example of the first table 8 in which the replacement process based on the second rule is executed (FIG. 13(b)).
[0121] As described above, the second rule is a rule for replacing (or converting) the designated terms (first description) of the first document with the description (second description) including their translated terms. As shown in FIG. 13(a), the file 7 including the second rule includes designated term information 71 for designating terms (designated terms) described in the first language, and translated term information 72 for designating the translated terms when the terms designated by the designated term information 71 are translated into the second language. As shown in FIG. 13(a), the designated term information 71 and the translated term information 72 each include one or more designated terms 71a and 71b and their translated terms 72a and 72b. The designated terms 71a and 71b designate the first description to be replaced.
[0122] <Processing Based on the Second Rule> In the activity A108 of FIG. 7, the replacement unit 113 executes a replacement process on the first table 6 shown in FIG. 12 based on the information included in the file 7. Thereby, a first table 8 as shown in FIG. 13(b) is obtained. In FIG. 13(b), columns 81a to 81d corresponding to columns 61a to 61d of the first table 6 in FIG. 12 are shown representatively.
[0123] The first elements 62b to 62d stored in columns 61b to 61d of FIG. 12 include the designated term 71a "information processing system" and the designated term 71b "processor" designated by the designated term information 71 in FIG. 13(a). These designated terms are replaced with the descriptions including the translated terms 72a "information processing system" and 72b "processor" in the first elements 82b to 82d stored in columns 81b to 81d of FIG. 13(b). Specifically, the description including the translated term includes the designated term and the translated term 72a or 72b with the identifier "<<>>" attached. That is, among the descriptions included in the first table 6 in FIG. 12, the designated term 71a "information processing system" is replaced with "information processing system<<information processing system>>", and the term 71b "processor" is replaced with "processor<<processor>>". On the other hand, since the first element 62a stored in column 61a of FIG. 12 does not include the designated term included in the designated term information 71, it remains as the first element 82a stored in column 81a of FIG. 13(b) as it is even after the substitution process based on the second rule is executed.
[0124] In this way, by executing the substitution process based on the second rule for the first table, it is possible to perform translation after substituting the description with the predetermined translation words, and it is easy to output the primary translation text as intended by the translation system. Also, it is possible to suppress the fluctuation of the translation words, that is, the conversion of the same term into different translation words.
[0125] In an environment where the information processing system 100 is used, when only the first document with a limited vocabulary is targeted for translation work, the second rule may also be stored in the storage unit 102, the storage unit 202, or the storage unit 302 in advance, similar to the first rule, and the substitution unit 113 may automatically read it via the reception unit 111. However, when translating the first document in a wide range of fields, it is preferable to let the user specify the second rule as in this embodiment. Thereby, it is possible to specify appropriate terms and their translation words according to the type of the first document. Also, compared with the case of applying a general-purpose second rule, since the capacity of the file 7 for specifying the second rule can be kept small, the load of the substitution process can be suppressed, contributing to shortening the time and reducing the load on the environment.
[0126] <Second File 9> In activities A109 to A115 of FIG. 8, the second table, which is the translation text based on the first table, is acquired, and formatting arrangements such as reverse substitution and insertion in a predetermined format are executed for the first table and the second table respectively, and the formatted first table and second table are integrated to generate the second file as shown in FIG. 14. FIG. 14 is a diagram showing an example of the second file 9. The second file 9 includes the first table 91 and the second table 93 arranged side by side.
[0127] As shown in FIG. 14, in the second file 9, each of a plurality of first elements 92 stored in the first table 91 and each of second elements 94 stored in the second table 93 are arranged so that their corresponding relationships can be understood. Specifically, to the right of the first elements 92a to 92j, the second elements 94a to 94d and 94e to 94j, which are their translated texts, are respectively arranged. In this way, by arranging the first element and the corresponding second element side by side, the user can discover the part to be edited while comparing the first element and the second element and perform operations such as editing.
[0128] In the second file 9, it is preferable that the horizontal width W91 of the first table 91 is less than or equal to half of the horizontal width W9 of the display area of the second file 9, and the horizontal width W93 of the second table 93 is determined according to the horizontal width W91 of the first table 91. According to such an aspect, when the first table and the second table are arranged, it can be within the horizontal width of the display area of the comparison file. That is, when referring to the second file 9, the first table 91 and the second table 93 can be viewed side by side left and right, and the editing work becomes easy.
[0129] Each of the plurality of first elements 92a to 92i respectively stored in the columns 91a to 91d, 91f to 91j of the first table 91 of the second file 9 is composed of the same text information as each of the first elements 52a to 52i of the first document 41 in FIG. 10 or the plurality of first elements 52 in FIG. 11. This is because in the activity A114 in FIG. 7, the first table 91 is obtained by performing an inverse substitution process on the plurality of first elements 82 stored in the first table 8 in FIG. 13(b). In other words, each of the plurality of first elements 92 included in the first table 91 is restored to the state of the original text.
[0130] Due to the reverse replacement in activity A114 of FIG. 7, for example, an empty column that does not store the first element may occur, such as column 91e of the first table 91 in FIG. 14. Therefore, the first table 91 itself is not exactly the same as the first table 5 in FIG. 11.
[0131] Each of columns 93a to 93j of the second table 93 of the second file 9 stores each of a plurality of second elements (primary translation texts) 94a to 94j. Among the text information included in the first elements 62g and 62j in FIG. 12 before the translation process is executed, the portions corresponding to "0001" and "0002" are half-width characters enclosed in angle brackets. These text information may remain in the original format and with the original angle brackets even after the translation process is executed. Therefore, in activity A113 of FIG. 7, for example, a process is executed to enclose "0001" and "0002" in corner brackets or convert them to full-width characters among the text information included in the second elements 94h and 94j. In this way, it can be replaced (or converted) into the format required by the Japan Patent Office where Japanese (the second language) is used. Note that in the first document, these paragraph numbers may be input using the automatic numbering function of word processing software instead of text information. Also in this case, by executing the process of activity A113, the paragraph numbers can be replaced with full-width numbers enclosed in corner brackets.
[0132] <The third file 17> By the processes of activities A122 to A124 in FIG. 9, a third file as shown in FIG. 15(a) is generated. FIG. 15 is a diagram showing an example of the third file 17 (FIG. 15(a)) and a diagram showing an example of the file 18 showing translation fee information (FIG. 15(b)).
[0133] As shown in FIG. 15(a), the third file 17 includes, in its display area, a second document (activity A123) generated by extracting each of the translated second elements. Further, in the third file 17, a formatting change process based on the fourth rule is being executed (activity A124). This formatting change process includes, for example, separating and arranging an item including text information of "claims" in the third file 17 and an item including text information of "specification" on separate pages 17a and 17b. Also, for example, it includes converting the text information that displays item names such as "claims" and "specification" to boldface. In this way, the fourth rule causes a formatting arrangement process to be executed to make the translated text easier to visually recognize. Note that when the first document and the second document are documents related to a patent or a utility model as in the present embodiment, the fourth rule may be a rule that causes a predetermined process to be performed so as to conform to the rules of the country where the second language is used, based on the second language.
[0134] <File 18 showing translation fee information> In activity A125 of FIG. 9, the translation fee is calculated and a file showing the translation fee information is generated. As shown in FIG. 15(b), in the file 18 showing the translation fee information, the number XX of drawings and the number YY of characters that are the calculation target of the translation fee are described as the basis of the translation fee. These pieces of information are obtained based on the first document as described above. Also, as the basis of the translation fee, the fee ZZ per word is also described. The information on this fee ZZ may be stored in the storage unit 102, the storage unit 202, or the storage unit 302, or may be input by the user each time. Also, for example, when the fee ZZ per word differs for each customer, reference information associating the customer information and the fee ZZ is stored, and the calculation unit 119 reads out the reference information via the reception unit 111 to obtain the fee ZZ per word and calculate the final amount to be billed WW. According to such an aspect, it is possible to present the translation fee calculated based on the number of characters or words in the original text. That is, the user does not need to manually calculate the billing amount.
[0135] According to this embodiment, by receiving a first file describing a first document which is the original text, it is possible to output a second file in which elements of the first document (first elements) and elements of the translated text (second elements) are associated and stored. Then, the user can appropriately edit the second file using a familiar word processing software or the like. Furthermore, based on the second file, it is possible to output a third file including a second document in which the format is adjusted and the elements of the translated text (second elements) are extracted, and a translation fee. In this way, according to the information processing system 100, it is possible to comprehensively support each process of the translation work from the editing of the translated text to the generation of the final deliverable.
[0136] [Modification Example] Hereinafter, a modification example of the above-described information processing system 100 will be described. The above-described embodiment and each of the following descriptions can be combined with each other.
[0137] In the above-described embodiment, the case where the server 10 and the translation system 20 are configured as separate entities has been described, but it is not limited thereto. For example, the translation system 20 may be included as a part of the server 10. That is, the storage unit 102 of the server 10 may store a translation program or the like stored in the storage unit 202 of the translation system 20, and the control unit 101 may execute the function as the control unit 201 of the translation system 20.
[0138] Furthermore, the user terminal 30 does not necessarily need to be configured as a separate entity. For example, the server 10 and the translation system 20 may be included as a part of the user terminal 30. That is, the storage unit 302 of the user terminal 30 may store programs or the like stored in the storage unit 102 of the server 10 and the storage unit 202 of the translation system 20, and the control unit 301 may execute the functions as the control unit 101 of the server 10 and the control unit 201 of the translation system 20.
[0139] In the above-described embodiment, the case where the user terminal 30 includes the display unit 304 and the control unit 301 (display control unit 313) has been described, but the present invention is not limited thereto. The user terminal 30 only needs to include a user interface. For example, the user terminal 30 may have an output device that appeals to the user's five senses, such as an audio output device or a haptic device, and the control unit 301 (display control unit 313) may cause the output device to express information such as sound or a sensory stimulus.
[0140] In the above-described embodiment, the case where the first rule, the third rule, and the fourth rule are stored in the storage unit 102 has been described, but the present invention is not limited thereto. At least one of the first, third, and fourth rules may be selected by the user, similarly to the second rule. Alternatively, the control unit 101 may be configured to present a plurality of options to the user and change the first, third, or fourth rule to be applied according to the user's selection result. The plurality of options may be, for example, the use of the third file, the customer to whom the third file should be delivered, and the like.
[0141] In the above-described embodiments, it has been described that each of the plurality of first elements included in the first table is obtained by dividing the first document into predetermined units such as one chapter, one page, one item, one paragraph, or one sentence. However, the present invention is not limited to this. For example, among the first elements, for a predetermined description or range, the unit of division may be configured to be changed. For example, when the first document includes a range or expression that should not be divided by a predetermined unit, it is preferable that reference information for excluding the range or expression from the division target in advance is stored in the storage unit 102 or the like. Alternatively, the reference information for dividing into the first elements may be changed and applied for each range such as an item. For example, when the first language is English and the unit of division is one sentence, it is preferable that the period attached to an abbreviation such as "FIG." is set not to be recognized as a sentence delimiter. Further, when the first document is a document related to a patent or a utility model, and the unit of division is one paragraph, it is preferable that in the claims or the claims for utility model registration, the line break is set not to be recognized as a paragraph delimiter. According to such an aspect, the first document can be divided into units that are easy for the user to recognize as one delimiter.
[0142] As described above, the order of each activity described with reference to the activity diagrams of FIGS. 7 to 9 can be appropriately changed, a plurality of processes may be executed simultaneously, or some processes may be omitted. For example, the inspection of the second table in activity A112 in FIG. 8 and the insertion of a predetermined format into the second table in activity A113 may be executed simultaneously. That is, while inspecting the second table, a predetermined format may be sequentially inserted.
[0143] Also, for example, the inspection in activity A112 and the insertion in activity A113 may be omitted. For example, when a specific format is not required for the third document which is a translated text, there will be no problem even if activities A112 and A113 are omitted.
[0144] In the above-described embodiment, in Activity A114 of FIG. 8, it has been described that it is preferable to restore the first table to the first table without replacement processing by inverse replacement, but the present invention is not limited thereto. For example, for a description corresponding to a description replaced based on at least the second rule, that is, a second description including both a designated term and a translated term with an identifier, the inverse replacement process may not be performed. Instead of the inverse replacement process, the control unit 101 may execute a process of deleting only the translated term with an identifier using the identifier as a mark, so as to return to a state where replacement by the second rule is not performed. In this way, for a universal rule (first rule) regardless of the first document, inverse replacement may be performed by referring to the rule again, and for a rule (second rule) specified differently by the first document, deletion using the identifier as a mark may be performed. Thereby, it is possible to balance the certainty of returning each of the first elements to the state of the original text and the lightness of the processing load.
[0145] Furthermore, neither inverse replacement nor deletion may be executed. Instead of these processes, the following may be executed. That is, the first document received in Activity A102 of FIG. 7 or the first table generated in Activity A103 may be stored in the storage unit 102 or the like. Then, when generating the second file in Activity A115 of FIG. 8, a first table for integration with the second table may be generated using the stored first document or first table.
[0146] In the above-described embodiment, as the third rule applied in Activity A119 of FIG. 9, it has been described that the claims or the utility model registration claims are repeated at a predetermined position, but the third rule is not limited thereto. For example, when the first document is a document related to a patent or a utility model, the third rule may be a rule for performing a predetermined process in accordance with the rules of the country where the third document, which is the translated text after translation, is submitted.
[0147] In the above-described embodiment, in Activity A125 of FIG. 9, when obtaining the number of drawings or tables, although it has been described that the arithmetic unit 119 obtains the number of drawings or tables based on the number of types of predetermined text information, it is not limited thereto. For example, when obtaining the number of drawings, the arithmetic unit 119 may obtain a predetermined item name as a mark. That is, for the first document included in the second file, text information such as "BRIEF DESCRIPTION OF DRAWINGS" or "Brief Explanation of Drawings" is searched, and the number of drawings may be obtained based on the number of lines changed in the item preceded by this text information. Further, the arithmetic unit 119 may obtain the number of drawings or tables based on the number of pieces of image information embedded in the first document.
[0148] Furthermore, the arithmetic unit 119 may also calculate the translation fee for the drawings or tables based on the number of words or characters. For example, the arithmetic unit 119 obtains the text information included in the drawing or table by OCR or the like, and multiplies the result of adding the number of words or characters included in the text information and the number of words or characters of the first document by one or more coefficients to obtain the final translation fee.
[0149] Also, although the case where the translation fee is calculated based on the number of characters or words of the first document has been described, it may be calculated based on other elements. For example, a part of the first document is changed to a predetermined format in advance, and the number of characters or words obtained by excluding the part in the predetermined format or multiplying the part in the format by a predetermined coefficient and increasing it is obtained, and the translation fee is calculated based on this.
[0150] Furthermore, for example, when the first document includes a fixed-form sentence and it is assumed that no translation fee is generated for the fixed-form sentence, the following may be done. That is, the third rule or the fourth rule applied in Activity A119 or Activity A124 may include executing a process of changing the fixed-form sentence to a predetermined format. Thereby, it is possible to automatically calculate and present the translation fee excluding the fixed-form sentence.
[0151] In each of the activities shown in FIGS. 7 to 9, each piece of information and each file output from the server 10 may simply be stored in the user terminal 30 or may be displayed on the display unit 304.
[0152] It may be provided in each of the aspects described below.
[0153] (1) A program that causes at least one computer to execute the following steps. In a first reception step, a first file including a first document described in a first language is received. In a first output step, a second file including a first table and a second table generated based on the first document is output. The first table divides the first document into a plurality of first elements and stores each of the plurality of first elements. The second table stores each of a plurality of second elements obtained by translating each of the plurality of first elements into a second language. The first table and the second table are arranged such that the correspondence between each of the plurality of first elements and each of the plurality of second elements can be understood. In a second reception step, the second file is received. In a second output step, a third file including a second document obtained by extracting the plurality of second elements from the second table included in the second file is output.
[0154] According to such an aspect, when a source file (first file) describing the source text (first document) is received, a comparison file (second file) comparing the elements of the source text (first elements) and the elements of the translated text (second elements) can be output. Also, a translated text file (third file) can be output from the comparison file (second file). Each of the first and second elements is, for example, a chapter, an item, a paragraph, or a sentence of a document, and is preferably a sentence. It is preset to exclude "FIG." etc.
[0155] (2) In the program described in (1) above, further, in the combining step, the first table and the second table are integrated into one table, where each of the plurality of corresponding first elements and each of the plurality of corresponding second elements are arranged side by side left and right for each corresponding element.
[0156] According to such an aspect, it is possible to output a comparison file in which each element of the original text and each element of the translated text are arranged side by side left and right for each corresponding element. Since it is easy to compare the elements of the original text and the translated text, it can satisfy the user.
[0157] (3) In the program described in (1) or (2) above, further, in the input step, the first table generated based on the first document is input into a translation system, and further, in the translation result receiving step, the second table obtained by translating the first table into the second language is received from the translation system.
[0158] According to such an aspect, it is possible to input the first table generated based on the original text into a translation system and obtain the translated second table.
[0159] (4) In the program described in (3) above, the width of the first table is less than or equal to half of the width of the display area of the second file, and the width of the second table is determined according to the width of the first table.
[0160] According to such an aspect, when the first table and the second table are arranged, it is possible to be within the width of the display area of the comparison file.
[0161] (5) In the program described in any one of (1) to (4) above, further, in the first replacement step, before translation, a process of replacing at least a part of the description of the first table is performed based on a predetermined first rule, and the first rule is a rule for pre-processing translation.
[0162] According to such an aspect, by replacing at least a part of the description of the original text based on the first rule before translation, a translation that better meets the user's requirements can be generated. For example, when the original text is a document related to a patent or a utility model, the first rule may correspond to the description specific to the patent or the utility model.
[0163] (6) In the program described in any one of (1) to (5) above, further, in the first replacement step, before performing the translation, a replacement process is performed on at least a part of the description of the first table based on a predetermined second rule. The second rule is a rule for replacing the first description of the first table with a second description. The first description is described in the first language, and the second description includes a description obtained by converting the first description into the second language.
[0164] In this way, by replacing with the second rule, it is possible to replace with a predetermined translation term and then perform the translation, making it easier for the translation system to output the intended translation. Also, it is possible to prevent the translation term from fluctuating.
[0165] (7) In the program described in (6) above, the second description includes both the first description and a description obtained by converting the first description into the second language and attaching an identifier.
[0166] With such a configuration, it is possible to restore the description of the first table to the original text simply by deleting the description in the second language.
[0167] (8) In the program described in any one of (5) to (7) above, further, in the second replacement step, after performing the translation, an inverse replacement process is performed on at least a part of the description of the first table corresponding to the replaced description.
[0168] Through such processing, specific descriptions such as Claims can be reliably restored to their original state. Note that for the descriptions replaced according to the second rule, it is not necessary to perform reverse replacement, or reverse replacement may be performed by referring to the second rule again.
[0169] (9) In the program according to any one of (5) to (8) above, in the first output step, the second file in which the first table that has not undergone replacement processing and the second table are arranged so as to be comparable for each of a plurality of elements is output.
[0170] According to such an aspect, the user can compare the translated text obtained by translating the original text after replacing it based on the first or second rule before translation with the original text.
[0171] (10) In the program according to any one of (1) to (9) above, further, in the execution step, based on each description of the plurality of first elements corresponding to each description of the plurality of second elements and the third rule, a predetermined process is executed for each description of the plurality of second elements, and the third rule includes a rule for formatting the second table.
[0172] According to such an aspect, in the comparison file, based on the description of the elements of the original text, a predetermined process can be performed on the elements of the translated text. Even if there are variations in the representation or descriptions in the elements of the translated text, a predetermined process can be performed based on the description of the original text. For example, when the original text is a document related to a patent or utility model, a predetermined process may be performed in accordance with the rules of the country where the translated text is to be submitted after translation.
[0173] (11) In the program according to any one of (1) to (10) above, the third file is a file obtained by executing a predetermined process for each description of the plurality of second elements based on the fourth rule, and the fourth rule includes a rule for changing the format of a predetermined description among the plurality of second elements.
[0174] According to such an aspect, it is possible to make it easy to read by making the item name bold or the like.
[0175] (12) In the program according to any one of (1) to (11) above, the first document is a document related to a patent or a utility model, and each of the plurality of first elements is obtained by dividing the first document in units of claims or paragraphs.
[0176] According to such an aspect, when the original text is a document related to a patent or a utility model, it is possible to compare the original text and the translated text for each claim or each paragraph. For example, it is convenient when considering correspondence for procedures abroad.
[0177] (13) In the program according to (12) above, further, in the insertion step, for each of the elements of the second table, when there is no claim number or paragraph number in a predetermined format, for each of the elements, the claim number or the paragraph number in the predetermined format is inserted.
[0178] According to such an aspect, when the original text is a document related to a patent or a utility model, it is possible to insert a claim number or a paragraph number in a predetermined format into the translated text.
[0179] (14) In the program according to any one of (1) to (13) above, further, in the presentation step, a translation fee calculated based on the number of characters or words of the first document is presented.
[0180] According to such an aspect, it is possible to present a translation fee calculated based on the number of characters or words of the original text.
[0181] (15) An information processing system comprising at least one processor, wherein the processor is configured to execute the program so that each step of the program according to any one of (1) to (14) above is performed.
[0182] According to such an aspect, when receiving a source file (the first file) that describes a source text (the first document), it is possible to output a comparison file (the second file) that compares the elements of the source text (the first elements) and the elements of the translated text (the second elements).
[0183] (16) An information processing method, comprising each step of the program described in any one of (1) to (14) above.
[0184] According to such an aspect, when receiving a source file (the first file) that describes a source text (the first document), it is possible to output a comparison file (the second file) that compares the elements of the source text (the first elements) and the elements of the translated text (the second elements). Of course, this is not the limit.
[0185] Finally, although various embodiments according to the present invention have been described, these are presented as examples and are not intended to limit the scope of the invention. The novel embodiments can be implemented in various other forms, and various omissions, replacements, and changes can be made without departing from the gist of the invention. The embodiments and their modifications are included in the scope and gist of the invention, and are also included in the invention described in the claims and the equivalent scope thereof.
Explanation of Reference Numerals
[0186] 100: Information processing system 10: Server 101: Control unit 111: Reception unit 112: Generation unit 113: Replacement unit 114: Inspection unit 115: Insertion unit 116: Integration unit 117: Execution unit 118: Extraction unit 119: Arithmetic unit 110: Output unit 102: Memory unit 103: Communication unit 104: Communication bus 17: Third file 17a: Page 17b: Page 18: File indicating translation fees 20: Translation system 201: Control unit 202: Memory unit 203: Communication unit 204: Communication bus 30: User terminal 301: Control unit 302: Memory unit 303: Communication unit 304: Display unit 305: Input unit 306: Communication bus 4: First file 41: First document 5: First table 51: Column 51a~g: Columns 52: First element 52a~g: First elements 6: First table 61: Column 61a~i: Columns 62: First element 62a~i: First elements 7: File 71: Specified term information 71a: Term 71b: Term 72: Translated word information 72a: Translated word 72b: Translated word 8: First table 81: Column 81a~d: Columns 82: First element 82a~d: First elements 9: Second file 91: First table 91a~j: Columns 92: First element 92a~j: First element 93: Second table 93a~j: Column 94: Second element 94a~j: Second element W9: Width W91: Width W93: Width WW: Claim amount ZZ: Fee
Claims
1. A program, causing at least one computer to execute the following steps: In a first reception step, receiving a first file including a first document described in a first language; In a first output step, outputting a second file including a first table and a second table generated based on the first document, wherein the first table divides the first document into a plurality of first elements and stores each of the plurality of first elements; the second table stores each of a plurality of second elements obtained by translating each of the plurality of first elements into a second language; the first table and the second table are arranged such that the correspondence between each of the plurality of first elements and each of the plurality of second elements is understandable; In a second reception step, receiving the second file; In a second output step, outputting a third file including a second document obtained by extracting the plurality of second elements from the second table included in the second file.
2. The program according to claim 1, further comprising a combining step of integrating the first table and the second table into one table, wherein each of the plurality of corresponding first elements and each of the plurality of second elements are arranged side by side left and right for each corresponding element.
3. The program according to claim 1, further comprising an input step of inputting the first table generated based on the first document into a translation system, and further comprising a translation result reception step of receiving, from the translation system, the second table obtained by translating the first table into the second language.
4. The program according to claim 3, wherein the width of the first table is at most half of the width of the display area of the second file, and the width of the second table is determined according to the width of the first table.
5. The program according to claim 1, further comprising a first replacement step of performing a replacement process on at least a part of the description of the first table based on a predetermined first rule before translation, wherein the first rule is a rule for preprocessing translation.
6. The program according to claim 1, Furthermore, in the first replacement step, before translation, a replacement process is performed on at least part of the descriptions in the first table based on a predetermined second rule. The second rule is a rule for replacing a first description in the first table with a second description. The first description is described in the first language. The second description includes a description obtained by converting the first description into the second language, and is a program.
7. In the program according to claim 6, the second description includes both the first description and a description obtained by converting the first description into the second language and attaching an identifier, and is a program.
8. In the program according to claim 5, Furthermore, in the second replacement step, after translation, a reverse replacement process is performed on at least part of the descriptions in the first table corresponding to the replaced description. Program.
9. In the program according to claim 5, In the first output step, a program that outputs the second file in which the first table that has not undergone the replacement process and the second table are arranged so as to be comparable for each of a plurality of elements.
10. In the program according to claim 1, Furthermore, in the execution step, based on the description of each of the plurality of first elements corresponding to the description of each of the plurality of second elements and a third rule, a predetermined process is executed on the description of each of the plurality of second elements. The third rule includes a rule for formatting the second table, and is a program.
11. In the program according to claim 1, The third file is a file obtained by executing a predetermined process on the description of each of the plurality of second elements based on a fourth rule. The fourth rule includes a rule for changing the format of a predetermined description among the plurality of second elements, and is a program.
12. In the program according to claim 1, The first document is a document related to a patent or a utility model. Each of the plurality of first elements is obtained by dividing the first document into claims or paragraphs, and is a program.
13. In the program according to claim 12, Furthermore, in the insertion step, for each of the elements of the second table, if there is no claim number or paragraph number in a predetermined format, a program for inserting the claim number or the paragraph number in the predetermined format for each of the elements.
14. In the program according to claim 1, Furthermore, in the presentation step, a program for presenting a translation fee calculated based on the number of characters or words of the first document.
15. An information processing system, comprising at least one processor, the processor is configured to execute the program such that each step of the program according to any one of claims 1 to 14 is performed.
16. An information processing method, a method comprising each step of the program according to any one of claims 1 to 14.
Citation Information
Patent Citations
Machine translation method, machine translation system and program
JP2016218995A