Japanese evaluation system, japanese evaluation apparatus, japanese evaluation method, and program
The Japanese language evaluation system addresses mistranslations by analyzing Japanese sentences for specific particle counts and issuing warnings, ensuring clarity and suitability for translation into foreign languages.
Patent Information
- Application Number
- JP2024103969
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-27
- Publication Date
- 2026-01-16
AI Technical Summary
Existing machine translation systems often produce mistranslations and unclear content due to differences in sentence structures between Japanese and foreign languages, despite pre-editing for undesirable words and phrases, as they do not account for the naturalness of sentences in Japanese.
A Japanese language evaluation system that performs morphological analysis to count the occurrences of specific particles like 'no' and sets detection counts before and after these particles, issuing warnings when the count exceeds a threshold to ensure clarity in Japanese sentences suitable for translation.
The system efficiently creates clear Japanese sentences by identifying and highlighting sections that may lead to mistranslations, enhancing the quality of machine translation into foreign languages.
Smart Images

Figure 2026005540000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a Japanese language evaluation system, a Japanese language evaluation device, a Japanese language evaluation method, and a program. [Background technology]
[0002] In recent years, with globalization, there has been an increasing demand for translating texts written in Japanese into foreign languages (especially English). This rise in demand has also led to an increase in demand for reducing the effort required for foreign language translation by using machine translation from Japanese to the foreign language. However, if post-editing is required after machine translation to correct mistranslations or unclear content, the effect of reducing the effort required is diminished.
[0003] The main reason for the need for post-editing is the Japanese original text before machine translation. Therefore, the problem can be solved to some extent by pre-editing the Japanese original text. For example, Patent Document 1 and Patent Document 2 describe document processing devices that issue a warning if a Japanese document contains undesirable words or phrases. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2000-090083 [Patent Document 2] Japanese Patent Application Laid-Open No. 2008-152349 Summary of the Invention [Problem to be solved by the invention]
[0005] However, because foreign languages and Japanese have different sentence structures, sentences that become unnatural when machine translated into a foreign language may appear natural in Japanese. Therefore, simply extracting undesirable words and phrases, as in Patent Documents 1 and 2, can result in mistranslations and unclear content. In other words, it is not possible to reduce the effort required for foreign language translation.
[0006] The present invention has been made in view of the above circumstances, and its purpose is to provide a Japanese language evaluation system, a Japanese language evaluation device, a Japanese language evaluation method, and a program that can efficiently create clear Japanese sentences that are also suitable for translation into foreign languages. [Means for solving the problem]
[0007] In order to solve the above problem, the invention described in claim 1 is a Japanese language evaluation system, a morphological analysis unit that divides the acquired Japanese target sentence into morphemes by morphological analysis; a counting unit that counts the number of times a predetermined particle is detected in a morpheme of the target sentence; and a setting unit that, when a predetermined morpheme is detected from the morphemes of the target sentence, divides the number of detections before and after the predetermined morpheme and sets the number of detections as the number of detections related to the target sentence.
[0008] The invention described in claim 2 is the Japanese language evaluation system described in claim 1, The apparatus further includes a determination unit that determines whether the number of detections set for the target sentence by the setting unit is equal to or greater than a predetermined value.
[0009] The invention described in claim 3 is the Japanese language evaluation system described in claim 2, The device further includes a warning unit that issues a warning when the determining unit determines that the number of detections is equal to or greater than the predetermined value.
[0010] The invention described in claim 4 is the Japanese language evaluation system described in claim 3, The warning unit highlights, in a document consisting of a plurality of target sentences, a target sentence whose detection count is equal to or greater than the predetermined value.
[0011] The invention described in claim 5 is the Japanese language evaluation system described in any one of claims 1 to 4, the measurement unit measures the number of times the predetermined particle is detected from the beginning of the target sentence, When the setting unit detects the specified morpheme between two of the specified particles, it sets the number of detections up to the specified morpheme as a first number of detections for the target sentence, and sets the number of detections after the specified morpheme as a second number of detections for the target sentence.
[0012] The invention described in claim 6 is the Japanese evaluation system described in any one of claims 1 to 4, The predetermined particle is "no".
[0013] The invention described in claim 7 is the Japanese evaluation system described in any one of claims 1 to 4, The predetermined morpheme is the particle "ga", the particle "wo", the particle "ni" or the particle "ya".
[0014] The invention described in claim 8 is the Japanese evaluation system described in any one of claims 1 to 4, The predetermined morpheme is a reading mark of an auxiliary symbol.
[0015] The invention described in claim 9 is the Japanese evaluation system described in any one of claims 1 to 4, When the target sentence includes a parenthetical expression, the measurement unit measures the number of times the predetermined particle is detected in the sentence within the parenthetical expression and in the sentence excluding the sentence within the parenthetical expression.
[0016] The invention described in claim 10 is a Japanese language evaluation device, an acquisition unit that acquires a Japanese target sentence; a morphological analysis unit that divides the acquired Japanese target sentence into morphemes by morphological analysis; a counting unit that counts the number of times a predetermined particle is detected in a morpheme of the target sentence; and a setting unit that, when a predetermined morpheme is detected from the morphemes of the target sentence, divides the number of detections before and after the predetermined morpheme and sets the number of detections as the number of detections related to the target sentence.
[0017] The invention described in claim 11 is A Japanese language evaluation method using a Japanese language evaluation device, a morphological analysis step of dividing the acquired Japanese target sentence into morphemes by morphological analysis; a measuring step of measuring the number of times a predetermined particle is detected in a morpheme of the target sentence; and a setting step of, when a predetermined morpheme is detected from the morphemes of the target sentence, dividing the number of detections before and after the predetermined morpheme and setting the number of detections as the number of detections related to the target sentence.
[0018] The invention described in claim 12 is a program, Computer, a morphological analysis unit that divides the acquired Japanese target sentence into morphemes through morphological analysis; a measurement unit that measures the number of times a predetermined particle is detected in a morpheme of the target sentence; When a predetermined morpheme is detected from the morphemes of the target sentence, the setting unit divides the number of detections before and after the predetermined morpheme and sets the number of detections as the number of detections related to the target sentence. [Effects of the Invention]
[0019] According to the present invention, clear Japanese sentences suitable for translation into foreign languages can be efficiently created. [Brief explanation of the drawings]
[0020] [Figure 1] FIG. 1 is a block diagram of a Japanese evaluation system. [Figure 2] 10 is a flowchart of a Japanese evaluation process. [Figure 3] 10 is a flowchart of a Japanese document analysis process. [Figure 4] FIG. 10 is a diagram showing an example of a manner in which one sentence in processing information is highlighted. DETAILED DESCRIPTION OF THE INVENTION
[0021] A Japanese language evaluation system including a Japanese language evaluation device according to an embodiment of the present invention will be described in detail below with reference to the accompanying drawings. However, the scope of the invention is not limited to the illustrated examples. In the following description, components having the same functions and configurations will be designated by the same reference numerals and their description will be omitted.
[0022] In the following, the Japanese document will be exemplified as a patent application specification, but the present invention is not limited to this.
[0023] [Japanese evaluation system] 1 is a block diagram showing the overall configuration of a Japanese language evaluation system 100. The Japanese language evaluation system 100 comprises a user terminal 10 and a Japanese language evaluation device 20. The user terminal 10 and the Japanese language evaluation device 20 are connected to each other via a network N.
[0024] The network N may include various communication networks such as telephone networks, optical fibers, mobile communication networks, communication satellite networks, CATV (Community Antenna Television) networks, and other dedicated lines, as well as internet service providers that connect each of these. The network N may also be a collective communication network in which various communication networks are interconnected so that they can communicate with each other. Examples of various communication networks include LANs (Local Area Networks), WANs (Wide Area Networks), WiFi (Wireless Fidelity), Bluetooth (registered trademark), and NFC (Near Field Communication). The connection method may be wired, wireless, or a combination of wired and wireless.
[0025] (user device) The user terminal 10 is a terminal device such as a desktop PC owned and used by each user of the Japanese language evaluation system 100. The user terminal 10 may also be a portable information terminal such as a smartphone, tablet terminal, or laptop computer.
[0026] Each user terminal 10 communicates data with the Japanese language evaluation device 20 using a network N. The user terminal 10 includes a control unit 11, a memory unit 12, a communication unit 13, an input unit 14, and a display unit 15. The various units of the user terminal 10 are connected by a bus 16.
[0027] {Control Unit} The control unit 11 is a processor (computer) that has a CPU (Central Processing Unit), ROM (Read Only Memory), RAM (Random Access Memory), etc., and controls the overall operation of each unit in the user terminal 10. The CPU reads out programs such as apps (application programs) stored in the storage unit 12, expands them into the RAM, and executes the programs to perform various processes.
[0028] {Storage section} The storage unit 12 includes a non-volatile semiconductor memory or the like, and stores programs such as apps and various data. A web browser or program capable of displaying processing information (described later) is installed on the user terminal 10. The files of the web browser or program are stored in the storage unit 12.
[0029] {Communications Department} The communication unit 13 is configured by a communication interface, etc. The communication unit 13 transmits and receives data to and from the Japanese language evaluation device 20 via the network N using a predetermined communication protocol.
[0030] {Input section} The input unit 14 receives an operation input from a user and transmits an operation signal corresponding to the operation input to the control unit 11. The input unit 14 includes an input device such as a keyboard, a mouse, or physical buttons.
[0031] {Display section} The display unit 15 includes a display device such as a liquid crystal display. The display unit 15 displays various information screens and the like under the control of the control unit 11. The input unit 14 and the display unit 15 may be integrated into a display device such as a liquid crystal panel and an input device such as a touch panel that is provided over the display screen of the display device.
[0032] 1, for the sake of simplicity, one user terminal 10 is shown as an example, but the present invention is not limited to this. A plurality of user terminals 10 may be connected to the Japanese language evaluation device 20 via a network N.
[0033] (Japanese evaluation device) The Japanese language evaluation device 20 is provided with the functions required in the Japanese language evaluation system 100. The Japanese language evaluation device 20 includes a control unit 21, a storage unit 22, and a communication unit 23. The various units of the Japanese language evaluation device 20 are connected by a bus 24.
[0034] The Japanese language evaluation device 20 is, for example, a server installed on the cloud. However, the Japanese language evaluation device 20 may also be installed as a physical device and operated on-premise. Furthermore, if the Japanese language evaluation device 20 is installed as a device, a display unit and an operation unit may be provided on the device so that various instructions can be input directly to the Japanese language evaluation device 20.
[0035] {Control Unit} The control unit 21 is composed of a CPU, RAM, etc. The CPU reads out various programs stored in the storage unit 22, expands them in the RAM, and executes various processes according to the expanded programs. The CPU centrally controls the operation of each unit in the Japanese language evaluation device 20. In particular, the control unit 21 functions as an acquisition unit 211, an analysis unit 212, a measurement unit 213, a setting unit 214, a determination unit 215, a generation unit 216, and an output unit 217 by the CPU executing the programs.
[0036] <Acquisition Department> The acquisition unit 211 acquires a document processing request together with a Japanese document from the user terminal 10 via the communication unit 23 .
[0037] <Analysis Department> The analysis unit 212 executes Japanese document analysis processing on the Japanese document acquired by the acquisition unit 211. The analysis unit 212 performs morphological analysis, syntactic analysis, etc. on the Japanese sentence using a program stored in the storage unit 22. In this way, the analysis unit 212 functions as a morphological analysis unit that divides the Japanese sentence into morphemes by morphological analysis.
[0038] <Measurement section> The measuring unit 213 measures the number of times a predetermined particle is detected in one sentence on which the analyzing unit 212 has performed the Japanese document analysis process.
[0039] In this embodiment, the predetermined particle is "no". This is because a sentence that uses the particle "no", which can have multiple meanings, multiple times can easily become ambiguous, redundant, or redundant when machine translated, making the meaning unclear.
[0040] <Settings Department> The setting unit 214 sets the number of times the particle "no" is detected for each sentence based on the measurement result of the measurement unit 213. For a sentence that includes a predetermined morpheme, the setting unit 214 sets the number of times the particle "no" is detected before and after the predetermined morpheme as a first number of times of detection and a second number of times of detection, respectively.
[0041] Specifically, the predetermined morpheme is the particle "ga," the particle "wo," the particle "ni," the particle "ya," or an auxiliary comma. As mentioned above, when the particle "no," which can have multiple meanings, is used multiple times in a sentence, the meaning of the sentence is likely to become unclear when machine translated. However, even when the particle "no" is used multiple times, if a predetermined morpheme exists in the middle of the multiple uses and the meaning of the sentence is cut off, the meaning of the sentence is unlikely to become unclear even when machine translated.
[0042] <Judgment section> The determination unit 215 determines whether the number of times the particle "no" is detected in each sentence, as measured by the measurement unit 213, is equal to or greater than a predetermined value. In this embodiment, the predetermined value is "3," but it may be greater than this. In addition, the predetermined value may be changeable as appropriate by an administrator of the Japanese language evaluation device 20, etc.
[0043] <Generation part> The generation unit 216 generates processing information as a result of the Japanese document analysis process by the analysis unit 212. The processing information is a file in a predetermined format that summarizes the parts and contents that need to be corrected or improved when machine translating a Japanese document that has been subjected to Japanese document analysis processing.
[0044] <Output section> The output unit 217 transmits the processing information generated by the generation unit 216 to the user terminal 10 via the communication unit 23. The processing information received by the user terminal 10 can be displayed on the display unit 15.
[0045] {Storage section} The storage unit 22 includes a hard disk drive (HDD), a nonvolatile semiconductor memory, etc. The storage unit 22 stores various programs and various data executed by the control unit 21. Note that at least a part of the various programs may be stored in the ROM of the control unit 21, etc.
[0046] In particular, the storage unit 22 stores natural language processing programs used by the analysis unit 212, such as a syntactic analysis program and a morphological analysis program. Examples of the syntactic analysis program include GiNZA, KNP, and CaboCha. Examples of the morphological analysis program include MeCab, JUMAN, Janome, Sudachi, and NLTK.
[0047] {Communications Department} The communication unit 23 is configured with a communication module or the like. The communication unit 23 transmits and receives various signals and various data to and from the user terminal 10 connected via the network N. The communication unit 23 is configured to receive data from an external device connected via the network N in a wired or wireless manner. Note that the communication unit 23 is not limited to being configured with a network interface or the like, and may be configured with a port or the like that can receive various data and various signals by inserting a USB memory, an SD card, or the like.
[0048] [Japanese evaluation processing] The Japanese language evaluation process performed by the Japanese language evaluation system 100 will be explained based on the flowcharts in Figures 2 and 3. The Japanese language evaluation system 100 executes the Japanese language evaluation process when the acquisition unit 211 receives a document processing request together with a Japanese language document from the communication unit 13 of the user terminal 10. As mentioned above, the following explanation will be given using a patent application specification as an example of a Japanese language document.
[0049] The analysis unit 212 divides the received patent application specification into sentence units (step S101). The analysis unit 212 can execute this process using a syntax analysis program stored in the storage unit 22. Alternatively, the analysis unit 212 can execute this process by detecting a period outside a parenthetical expression.
[0050] The analysis unit 212 detects paragraphs in the received patent application specification (step S102). Since patent application specifications have a fixed format, a section from a predetermined fixed phrase to another predetermined fixed phrase is detected as one paragraph. For example, in a patent application specification, the section from "Background Art" to the sentence immediately preceding "Prior Art Literature" is detected as the "Background Art" paragraph.
[0051] The analysis unit 212 performs Japanese document analysis processing on the sentence of the paragraph to be processed among the divided sentences (step S103). A detailed flowchart of the Japanese document analysis processing is shown in FIG.
[0052] The paragraphs to be processed in a patent application specification are "Technical Field," "Background Art," "Problem to be Solved by the Invention," "Brief Description of the Drawings," "Form for Carrying Out the Invention," "Industrial Applicability," "Abstract," and "Claims."
[0053] The analysis unit 212 performs parenthetical processing (step S103a). Specifically, for example, assume that the paragraph to be processed contains a sentence, "Method for generating a color image (image of ink of multiple colors)." In this case, the analysis unit 212 divides the sentence into two independent sentences, "Method for generating a color image" and "Image of ink of multiple colors."
[0054] The analysis unit 212 performs morphological analysis (step S103b). The analysis unit 212 divides each divided sentence in the paragraph to be processed into morphemes, which are the smallest units in language, through the morphological analysis.
[0055] The measurement unit 213 extracts the particle "no" in one sentence in the paragraph to be processed, starting from the beginning of the sentence (step S103c). Then, the measurement unit 213 determines whether or not the particle "no" has been detected in the one sentence (step S103d).
[0056] In particular, if the morpheme immediately preceding the particle "no" is a noun, an adjective, the particle "nado", an auxiliary symbol, or a suffix, the measuring unit 213 recognizes the particle "no" as a detection target.
[0057] If the measuring unit 213 does not detect the particle “no” (step S103d; No), the setting unit 214 sets the number of times it has been detected in the sentence to “0” (step S103e), and the process proceeds to step S103k, which will be described later.
[0058] If the particle "no" is detected (step S103d; Yes), the measurement unit 213 sets the number of times the particle "no" has been measured to 1 (step S103f). Then, the measurement unit 213 determines whether or not the particle "no" has been further detected in the one sentence (step S103g).
[0059] If the particle "no" is further detected (step S103g; Yes), the measurement unit 213 determines whether or not a predetermined morpheme exists between the two particles "no" (step S103h).
[0060] If a predetermined morpheme exists between the two particles "no" (step S103h; Yes), the setting unit 214 sets the measured number of times the particle "no" has been used up to the predetermined morpheme in the sentence as the nth detection count (n: loop count) (step S103i). Then, the measurement unit 213 proceeds to step S103f, resets the measured number of times the particle "no" has been used to 1, and continues detecting the particle "no". If a predetermined morpheme does not exist between the two particles "no" (step S103h; No), the measurement unit 213 increments the measured number of times by 1 (step S103j). Then, the process proceeds to step S103g, and continues detecting the particle "no".
[0061] If the particle "no" is not further detected (step S103g; No), in other words, if the detection of the particle "no" in the one sentence is completed, the setting unit 214 sets the number of times of detection in the one sentence (step S103k). Note that, if multiple detection numbers are set in step S103i, the multiple detection numbers are set as the number of times of detection in the one sentence.
[0062] Furthermore, in step S103a, if a sentence containing parentheses is divided into separate sentences, that is, the sentence inside the parentheses and the sentence excluding the parenthesized portion, the setting unit 214 sets the number of times the particle "no" is detected in the sentence inside the parentheses separately and independently from the sentence outside the parentheses.
[0063] For example, suppose there is a sentence that reads, "Method for generating a color image (image of ink of multiple colors)." In this case, the measurement unit 213 counts the particle "no" twice in the sentence outside the parentheses, "Method for generating a color image." Then, the measurement unit 213 counts the particle "no" three times in the sentence inside the parentheses, "image of ink of multiple colors." Therefore, the setting unit 214 sets the number of detections in the example sentence to "2" and "3," respectively.
[0064] The measurement unit 213 determines whether the paragraph to be processed contains a sentence for which the detection count has not been set, i.e., a sentence for which the particle "no" has not yet been extracted (step S103l). If a sentence for which the detection count has not been set exists (step S103l; Yes), the process proceeds to step S103c, where extraction of the particle "no" continues. If there is no sentence for which the detection count has not been set (step S103l; No), the Japanese document analysis process ends.
[0065] After the Japanese document analysis process is completed, the determination unit 215 performs a determination process (step S104). The determination unit 215 determines whether the number of times the particle “no” is detected is equal to or greater than a predetermined value (for example, 3) for each sentence in the paragraph to be processed.
[0066] The generating unit 216 generates processing information (step S105). The processing information is, for example, a warning added to the patent application specification acquired by the acquiring unit 211 by highlighting a sentence in which it is determined that the number of times the particle "no" is detected is equal to or greater than a predetermined value as a result of the determination process. In this way, the generating unit 216 functions as a warning unit that notifies the user of a warning when it is determined that the number of times the particle is detected is equal to or greater than a predetermined value.
[0067] The output unit 217 performs an output process of transmitting the processing information generated by the generation unit 216 to the user terminal 10 (step S106).
[0068] The processing from step S103c to step S106 will be explained below based on a specific example. In this example, the explanation will be based on the sentence, "This makes it possible to make the transport operation of some of the valid sheets different from the transport operation of other sheets."
[0069] The measurement unit 213 proceeds to extract the particle "no" from the beginning of the sentence (step S103c). The measurement unit 213 first detects that the "no" immediately after the noun "kakkou gami" is a particle (step S103d; Yes), and sets the measurement count of the particle "no" to "1" (step S103f). Next, the measurement unit 213 detects that the "no" immediately after the noun "kabu" is a particle (step S103g; Yes). At this time, between the particle "no" immediately after "kakkou gami" and the particle "no" immediately after "kabu", only the noun "kabu" exists as a morpheme, and no predetermined morpheme exists (step S103h; No). Therefore, the measurement unit 213 increments the measurement count by 1 (step S103j).
[0070] Next, the measurement unit 213 detects that the "no" immediately after the noun "transport" is a particle (step S103g; Yes). At this time, between the particle "no" immediately after "bukata" and the particle "no" immediately after "transport", the only morpheme present is the noun "transport", and no predetermined morpheme exists (step S103h; No). Therefore, the measurement unit 213 adds 1 to the number of measurements (step S103j).
[0071] Next, the measurement unit 213 detects that "no" immediately after the noun "ta" is a particle (step S103g; Yes). At this time, between the particle "no" immediately after "touyou" and the particle "no" immediately after "ta", there is a particle "wo" which is a predetermined morpheme immediately after the noun "doushi" (operation) (step S103h; Yes). Therefore, the setting unit 214 sets the measurement number up to the particle "wo" to "3" as the first detection number (step S103i). Then, the measurement unit 213 resets the measurement number to 1 (step S103f).
[0072] Next, the measurement unit 213 detects that the particle "no" immediately after the noun "paper" is a particle (step S103g; Yes). At this time, the only morpheme between the particle "no" immediately after "other" and the particle "no" immediately after "paper" is the noun "paper," and no predetermined morpheme exists (step S103h; No). Therefore, the measurement unit 213 adds 1 to the number of measurements (step S103j).
[0073] There is no particle "no" after "no" immediately after "paper" (step S103g; No). Therefore, the setting unit 214 sets the number of measurements after the particle "o" to "2" as the second detection count. Then, the setting unit 214 sets the first detection count and the second detection count to "3" and "2" as the detection counts of the sentence (step S103k).
[0074] Step S103l will be omitted. Because the first detection count for the sentence is "3," the determination unit 215 determines that the number of detections of the particle "の" is equal to or greater than a predetermined value (step S104). Therefore, the generation unit 216 generates processing information in which the sentence is highlighted, as shown in FIG. 4 (step S105). Then, the output unit 217 transmits the processing information to the user terminal 10 (step S106). Note that highlighting refers to a state in which a specific sentence or a specific portion of a sentence is colored, enclosed, underlined, or the like, or the font is changed so that the specific sentence or specific portion can be distinguished or identified from the rest. Furthermore, the sentence or specific portion to be highlighted may be a single sentence or a mixture of multiple sentences.
[0075] [Effects of the embodiment] As described above, the analysis unit 212 of the Japanese language evaluation system 100 divides the acquired Japanese document into morphemes. The measurement unit 213 measures the number of times a predetermined particle is detected in the morphemes of the target sentence. When the setting unit 214 detects a predetermined morpheme from the target sentence, it sets the first and second detection counts before and after the predetermined morpheme as the detection count for the target sentence. With this configuration, the presence of a predetermined morpheme that cuts off the meaning of the sentence can be reflected in the detection count of the predetermined particle, making it possible to efficiently create clear Japanese sentences suitable for foreign language translation.
[0076] [Other configurations] Although the present invention has been specifically described above based on the embodiments thereof, the present invention is not limited to the above-described embodiments. Of course, the present invention can be modified in various ways within the scope of the invention described in the claims and its equivalents.
[0077] For example, in the above example, the measuring unit 213 extracts the particle "no" in one sentence in order from the beginning of the sentence, measures the number of times it is detected, and when a predetermined morpheme is detected, the setting unit 214 sets the first and second detection numbers, but this is not limiting. That is, the measuring unit 213 may detect all particles "no" in one sentence, and then detect the predetermined morpheme, and the setting unit 214 may set the number of times the particle "no" is detected in each section divided by the predetermined morpheme as the first and second detection numbers, respectively.
[0078] In the above, the particle "ga," the particle "wo," the particle "ni," the particle "ya," or the auxiliary comma are given as examples of the predetermined morpheme, but the present invention is not limited to these. Verbs between the particle "no," auxiliary verbs, conjunctions, particles other than "no," and particles other than "nado" may also be the predetermined morpheme.
[0079] Furthermore, in the above description, the generation unit 216 highlights sentences in which the number of times the particle "no" is detected is three or more, but this is not limited to this. That is, the particle to be determined by measuring the number of times it is detected in one sentence is not limited to "no". For example, the particle "de", which can have multiple meanings, may be added and / or switched as a particle to be detected.
[0080] Also, while Figure 4 shows an example in which the entire sentence where the number of detections is equal to or greater than a predetermined value is highlighted, this is not limited to the case where a predetermined morpheme is included in the sentence. If the number of detections is equal to or greater than a predetermined value only for a portion before or after the predetermined morpheme, such as the portion up to "This causes..." in the sentence in Figure 4, only that portion may be highlighted. This configuration makes it clearer which parts need to be corrected, enabling the efficient creation of clear Japanese sentences suitable for foreign language translation.
[0081] Furthermore, the processing content of the generation unit 216 is not limited to issuing a warning by highlighting a sentence in which the number of times the particle "no" is detected is a predetermined value or more. For example, a warning may be issued by highlighting a sentence with a predetermined number of characters or more, a sentence determined by a syntax analysis program to have multiple subjects, or a sentence containing a predetermined word or expression that is undesirable to use. Note that it is preferable to appropriately differentiate the highlighting style of these sentences from the highlighting style of a sentence in which the number of times the particle "no" is detected is a predetermined value or more, so that the content of the warning can be distinguished.
[0082] Furthermore, in the above example, the Japanese document is a patent application specification, but this is not limited to this, and any document written in Japanese can be applied as a processing target for the Japanese language evaluation system 100 of the present invention.
[0083] Furthermore, although hard disks and semiconductor nonvolatile memories have been exemplified above as computer-readable media for the program according to the present invention, the present invention is not limited to these examples. Other computer-readable media include portable recording media such as CD-ROMs. Furthermore, carrier waves are also applicable as a medium for providing data for the program according to the present invention via a communication line. [Explanation of symbols]
[0084] 100 Japanese rating system 20 Japanese evaluation device 211 Acquisition Department 212 Analysis Department (Morphological Analysis Department) 213 Measurement Department 214 Settings 215 Judgment section 216 Generation part (warning part)
Claims
1. a morphological analysis unit that divides the acquired Japanese target sentence into morphemes by morphological analysis; a counting unit that counts the number of times a predetermined particle is detected in a morpheme of the target sentence; A Japanese evaluation system comprising: a setting unit that, when a specified morpheme is detected from the morphemes of the target sentence, divides the number of detections before and after the specified morpheme and sets them as the number of detections for the target sentence.
2. The Japanese language evaluation system according to claim 1 , further comprising a determination unit that determines whether the number of detections set for the target sentence by the setting unit is equal to or greater than a predetermined value.
3. The Japanese language evaluation system according to claim 2 , further comprising a warning unit that issues a warning when the determining unit determines that the number of detections is equal to or greater than the predetermined value.
4. 4. The Japanese language evaluation system according to claim 3, wherein the warning unit highlights a target sentence whose detection count is equal to or greater than the predetermined value in a document consisting of a plurality of target sentences.
5. the measurement unit measures the number of times the predetermined particle is detected from the beginning of the target sentence, 5. The Japanese evaluation system according to claim 1, wherein, when the specified morpheme is detected between two of the specified particles, the setting unit sets the number of detections up to the specified morpheme as a first number of detections for the target sentence, and sets the number of detections after the specified morpheme as a second number of detections for the target sentence.
6. The Japanese evaluation system according to claim 1 , wherein the predetermined particle is “no”.
7. The Japanese evaluation system according to claim 1 , wherein the predetermined morpheme is a particle "ga", a particle "wo", a particle "ni", or a particle "ya".
8. The Japanese evaluation system according to claim 1 , wherein the predetermined morpheme is a comma, which is an auxiliary symbol.
9. 5. The Japanese evaluation system according to claim 1, wherein, when the target sentence includes a parenthetical expression, the measurement unit measures the number of times the specified particle is detected in the sentence within the parenthetical expression and in the sentence excluding the sentence within the parenthetical expression.
10. An acquisition part that acquires Japanese target sentences, a morphological analysis unit that divides the acquired Japanese target sentence into morphemes by morphological analysis; a counting unit that counts the number of times a predetermined particle is detected in a morpheme of the target sentence; a setting unit that, when a predetermined morpheme is detected from the morphemes of the target sentence, divides the number of detections before and after the predetermined morpheme and sets the number of detections as the number of detections related to the target sentence.
11. A Japanese language evaluation method using a Japanese language evaluation device, a morphological analysis step of dividing the acquired Japanese target sentence into morphemes by morphological analysis; a measuring step of measuring the number of times a predetermined particle is detected in a morpheme of the target sentence; A Japanese evaluation method comprising: a setting step in which, when a specified morpheme is detected from the morphemes of the target sentence, the number of detections is divided into those before and after the specified morpheme and set as the number of detections for the target sentence.
12. Computer, a morphological analysis unit that divides the acquired Japanese target sentence into morphemes through morphological analysis; a measurement unit that measures the number of times a predetermined particle is detected in a morpheme of the target sentence; A program that functions as a setting unit that, when a specified morpheme is detected from the morphemes of the target sentence, divides the number of detections before and after the specified morpheme and sets them as the number of detections for the target sentence.
Citation Information
Patent Citations
Document processor and detailed statement processor
JP2000090083A
Document data processor
JP2008152349A