Language processing computer, language processing method and program

The language processing computer improves AI translation accuracy by analyzing and translating documents in semantic units, addressing cultural and contextual nuances to enhance precision in fields like law, finance, and medicine.

JP7811816B1Active Publication Date: 2026-02-06NISEKO UNLEASHED LLC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025156521
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2026-02-06
Estimated Expiration
2045-09-19

AI Technical Summary

Technical Problem

Conventional AI learning models fail to capture language-specific cognitive structures and emotional nuances, leading to misreadings of context and intent due to cultural differences and mistranslations, especially in fields like law, finance, and medicine.

Method used

A language processing computer that performs syntactic analysis, maps words to semantic structures, evaluates similarity and relationships, and extracts character strings in units of phrases, propositions, or logic using a semantic unit dictionary to improve analysis and translation accuracy.

Benefits of technology

Enhances language analysis and translation accuracy by capturing semantic units, reducing mistranslations, and ensuring logical consistency, particularly in fields where precision is critical.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007811816000001_ABST
    Figure 0007811816000001_ABST
Patent Text Reader

Abstract

A language processing computer is provided that can improve the accuracy of language analysis and the accuracy of semantic analysis, which are the prerequisites for learning by artificial intelligence. [Solution] The language processing computer 1 comprises a syntactic analysis unit 110 that performs syntactic analysis of an input document and divides the document into word-based elements; a mapping unit 120 that maps each of the divided words to a semantic structure; an evaluation unit 130 that evaluates the semantic match or similarity between the words divided into word-based elements; a memory unit 300 that associates each word, the evaluated score, and the document and stores them as a semantic unit dictionary; an extraction unit 210 that extracts character strings from a specified document in units of phrases, propositions, or logic based on the semantic unit dictionary; and a processing unit 220 that processes the analysis or translation of the specified document based on the extracted character strings.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a language processing computer, a language processing method, and a program for performing language processing, which is a prerequisite for training an artificial intelligence. [Background technology]

[0002] Conventional AI (Artificial Intelligence) learning models, when learning multiple languages, rely on direct one-to-one translation of words and documents, often failing to capture differences in language-specific cognitive structures and emotional nuances. As a result, AI learning is subject to misreading of context and intent, and response errors due to cultural differences (mistranslations and failure to reflect subtle differences in meaning). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2025-60552 [Non-Patent Document 1] Internet "AI Compass," "Tokenization for Understanding Language Grain," What is Tokenization?, published February 1, 2025, retrieved August 13, 2025, Internet URL <https: / / ai-compass.weeybrid.co.jp / llm / understanding-granular-language-tokenization / #:~:text=%E4%BA%BA%E9%96%93%E3%81%AF%E6%96%87%E7%AB%A0%E3%82%92%E8%AA%AD%E3%82%80,%E3%81%A7%E3%81%82%E3%81%A3%E3%81%9F%E3%82%8A%E3%82%82%E3%81%97%E3%81%BE%E3%81%99%E3%80%82> Summary of the Invention [Problem to be solved by the invention]

[0004] For example, Patent Document 1 discloses that when searching for information, the language is translated and the search is performed, and that in this case, a predefined emotion dictionary is used to analyze emotions, translate text characters, and perform information searches, etc.

[0005] However, the emotion dictionary that is referenced is merely a database that lists words that express emotions, and the accuracy of information retrieval that reflects emotions cannot be expected. Furthermore, as in Non-Patent Document 1, language analysis is performed by forming units called tokens, which are strings of multiple characters obtained by dividing the components of a sentence into words. For example, existing generative AI defines token units as strings of characters, which leads to problems such as a mismatch between breaks in meaning and token units, semantic units collapsing during translation, and the core of the sentence structure (subject, predicate, relative clause) becoming disintegrated, and this type of language analysis does not necessarily capture the meaning of a sentence appropriately.

[0006] Therefore, the object of this invention is to improve the accuracy of language analysis by providing a processing method based on "semantic units," which are the smallest unit that conveys meaning and are groups of comparable abstraction. In other words, we provide a language processing computer, language processing method, and program that improve the accuracy of language analysis and semantic analysis, which are the prerequisites for learning by artificial intelligence. [Means for solving the problem]

[0007] The present invention includes a syntactic analysis unit that performs syntactic analysis of an input document and divides the document into elements on a word-by-word basis; a mapping unit that maps each of the decomposed words to a semantic structure; Evaluating the semantic agreement or similarity between the words divided into the word-unit elements Generate a similarity score by an evaluation unit; Each of the words and Generatea memory unit for storing a semantic unit dictionary as a database in which the scores are associated with the documents, thereby mapping each word to a semantic structure and indicating the similarity between the words, and relationship tags that indicate the relationship between the words or propositions and include causation, inclusion, contrast, subject-predicate-object relationship, and superordinate-subordinate relationship; an extraction unit that extracts character strings from a predetermined document in units of phrases, propositions, or logic based on the semantic unit dictionary; a processing unit that processes analysis or translation of the predetermined document based on the extracted character string; A language processing computer is provided.

[0008] According to the present invention, words are divided into words with meaning, and the agreement or similarity between the words is evaluated by a score, and the words are stored as a semantic unit dictionary. Analysis or translation is performed based on character strings extracted from the semantic unit dictionary in units of phrases, propositions, or logic, thereby improving the accuracy of the analysis or translation.

[0009] In addition, in the present invention, the semantic unit dictionary may include the relationship tags that indicate the relationships between the words divided into the word-unit elements, and a set of reference propositions that serve as standards for determining consistency or contradiction with the propositions in the specified document, which are a set of multiple propositions that are registered in advance to ensure logical accuracy.

[0010] In addition, in the present invention, the extraction unit may refer to the score, the group of reference propositions, and the relation tag contained in the semantic unit dictionary, and extract the character string after classifying it into one of the phrase, the proposition, or the logic.

[0011] Further, the present invention is a method for extracting a plurality of data from a plurality of data. extracting the character string related to the phrase by determining the phrase based on the score included in the semantic unit dictionary; extracting the character strings related to the proposition by comparing the logical structure of the subject, predicate, and object of the proposition with the reference propositions included in the semantic unit dictionary and evaluating whether they match or not; For the logic, the character strings related to the proposition may be extracted by determining the relationship between the character strings related to the logic based on the relation tags included in the semantic unit dictionary.

[0012] This allows strings related to phrases to be extracted based on the evaluated score, strings related to propositions to be extracted based on a group of reference propositions, and strings related to logic to be extracted based on relational tags, making it possible to correctly extract strings in units of phrases, propositions, or logic.

[0013] Furthermore, the present invention provides, as basic data, character strings extracted from the predetermined document in units of the phrase, the proposition, or the logic based on the score, the standard proposition group, or the relation tag contained in the semantic unit dictionary, with labels indicating possible analysis or translation errors added; The processing unit may correct analysis accuracy or translation accuracy using the basic data.

[0014] According to this method, a label indicating a possible analysis or translation error is added to the extracted character string, and the label can be used to correct the analysis accuracy or translation accuracy.

[0015] The present invention basically belongs to the category of things as a language processing computer, but the same effects or advantages can be achieved by a language processing method executed by a language processing computer and a program readable by a language processing computer. [Effects of the Invention]

[0016] According to the present invention, it is possible to improve the accuracy of language analysis, and to improve the accuracy of language analysis and semantic analysis, which are the prerequisites for AI learning. For example, in fields such as law, finance, medicine, and education, where mistranslation is unacceptable, it is possible to obtain higher evaluation results than existing AI for contracts, medical records, etc., based on indicators such as logical consistency, nuance preservation rate, and mistranslation reduction rate. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a diagram showing an overview of a language processing computer according to an embodiment of the present invention. [Figure 2] 1 is a block diagram showing the functional configuration of a language processing computer according to an embodiment of the present invention. [Figure 3] FIG. 2 is a diagram schematically illustrating a semantic unit dictionary (main body) in a language processing computer according to one embodiment of the present invention. [Figure 4] FIG. 1 is a diagram schematically showing similar concept clusters and a group of reference propositions in a language processing computer according to an embodiment of the present invention. [Figure 5] 10 is a flowchart showing the flow of a dictionary creation and update process performed by a language processing computer according to one embodiment of the present invention. [Figure 6] 1 is a flowchart showing the flow of a document analysis and translation process performed by a language processing computer according to an embodiment of the present invention. [Figure 7] 1 is a flowchart showing the flow of a conventional token-based document analysis and translation process. DETAILED DESCRIPTION OF THE INVENTION

[0018] Hereinafter, a mode for carrying out the present invention (hereinafter, referred to as an embodiment) will be described in detail with reference to the accompanying drawings. However, the embodiment described below is merely an example of the present invention, and the present invention is not limited thereto. Furthermore, the effects described in the embodiment of the present invention are merely a list of the most preferable effects resulting from the present invention, and the effects of the present invention are not limited to those described in the embodiment of the present invention.

[0019] [Outline of language processing computers and related devices] The language processing computer 1 according to this embodiment creates and updates a dictionary organized by semantic units (semantic unit dictionary), analyzes and translates input documents through unique processing at three levels—phrases, propositions, and logic—described below, and outputs text that matches or approximates natural language through linguistic and semantic analysis of the document. A semantic unit is the smallest unit that conveys meaning and is different from a word-based token. The language processing computer 1 is suitable for use in fields where mistranslation is unacceptable, such as law, finance, medicine, and education. In the following description, a document is considered to mean textual or audio information consisting of one or more words that have meaning as a phrase or sentence. A word is also considered to include a phrase consisting of multiple words.

[0020] As shown in FIG. 1, a language processing computer 1 is communicably connected via a network (Internet) to an AI server 2 having a language knowledge base and a user terminal 3 such as a user's personal computer or smartphone that inputs documents.

[0021] The language processing computer 1 includes a processor 10, a mass storage device 11, an input device 12, an output device 13, and a communication device 14 as basic components.

[0022] The processor 10 is a device that mainly includes a CPU (Central Processing Unit) that interprets instructions given by a program, performs various calculations, and executes various controls, and also includes a ROM (Read Only Memory) and a RAM (Random Access Memory) as main storage devices.

[0023] The mass storage device 11 is an auxiliary storage device with a relatively large storage capacity, such as a hard disk drive or solid state drive, and stores therein programs and data to be executed by the processor 10, as well as a semantic unit dictionary (described later) in a rewritable manner.

[0024] The input device 12 is a device such as a keyboard, mouse, microphone, camera, etc., for inputting information to the language processing computer 1, and includes an input interface for taking in information via the communication device .

[0025] The output device 13 is a device such as a display, printer, speaker, etc. that outputs the processed results and information, and includes an output interface that sends information via the communication device 14.

[0026] The communication device 14 is a device for transmitting and receiving information via a network (Internet).

[0027] The AI ​​server 2 has a core configuration for artificial intelligence, and is capable of performing AI processing other than the functions possessed by the language processing computer 1, and also has a large-scale language knowledge base such as WordNet.

[0028] The user terminal 3 is a device such as a personal computer or smartphone that allows a user to input a document as text information or voice information and instruct the language processing computer 1 to analyze or translate the input document via a network (Internet). Note that the language processing computer 1 may also be configured so that a user directly inputs a document and instructs the language processing computer 1 to analyze or translate the input document. In other words, the user terminal 3 may be configured to input and output documents using, for example, a chatbot or a voice assistant, and may have the same functional configuration as the language processing computer 1.

[0029] [Functional configuration of language processing computers] 2, the language processing computer 1 includes, as functional components related to language processing, an input unit 100, a syntax analysis unit 110, a mapping unit 120, an evaluation unit 130, a determination unit 140, an automatic generation unit 150, an output unit 200, an extraction unit 210, a processing unit 220, a re-evaluation unit 230, and a storage unit 300. The input unit 100 and the output unit 200 are realized by an input device 12 and an output device 13. The storage unit 300 is realized by a large-capacity storage device 11. The syntax analysis unit 110, the mapping unit 120, the evaluation unit 130, the determination unit 140, the automatic generation unit 150, the extraction unit 210, the processing unit 220, and the re-evaluation unit 230 are realized by the processor 10 executing a program.

[0030] The input unit 100 accepts documents (hereinafter referred to as user documents) input from the user terminal 3. The input unit 100 also accepts documents (hereinafter referred to as "public documents") that have been made public on the network by so-called bot processing.

[0031] The parsing unit 110 performs a syntactic analysis of documents input as user documents or public documents, and breaks the documents down into elements on a word-by-word basis. For example, if a document such as "The spread of electrically assisted bicycles is changing urban mobility," is input, the parsing unit 110 breaks the document down into words such as "electrically assisted bicycles," "popularity," "urban mobility," and "changing" through syntactic analysis of subjects, predicates, objects, conjunctions, etc.

[0032] The mapping unit 120 maps each word decomposed by the parser 110 to a semantic structure. The semantic structure is a semantic space (semantic space) in which word meanings are expressed and arranged multidimensionally. Differences in meaning are expressed vectorially, and the meanings converted into numerical vectors are arranged in space. For example, words such as "electrically assisted bicycle," "popular," "urban mobility," and "changing" are mapped to words with meanings such as "bicycle," "increase," "urban mobility," and "change" arranged in the semantic structure (semantic space). The mapped words are the smallest units of meaning, i.e., semantic units. Furthermore, the mapping unit 120 acquires, as meta-information for each word, location information indicating the location of the word in the document and context information indicating the context surrounding the location of the word. The meta-information is associated with each word and registered in the semantic unit dictionary (main body) described below.

[0033] The evaluation unit 130 evaluates the semantic match or similarity between words mapped to the semantic structure by the mapping unit 120. The semantic match or similarity is calculated as a numerical score (similarity score) based on at least one of the following parameters: (1) conceptual distance in the semantic structure (semantic space): for example, the distance measured in the semantic space of "electrically assisted bicycle" and "popular" (a space in which things with similar meanings are closer and things with different meanings are farther apart); (2) co-occurrence frequency: for example, a numerical value indicating how often "electrically assisted bicycle" and "popular" appear together; (3) cosine similarity between embedding vectors; and (4) the match rate of definitions and descriptions. The similarity score is associated with each word in the decomposed semantic unit and registered in the semantic unit dictionary (main body), which will be described later.

[0034] The evaluation unit 130 also evaluates the relationships between words mapped to the semantic structure by the mapping unit 120, and generates relational tags as the evaluation results. Relational tags are labels that indicate the semantic and structural relationships between words (such as superordinate and subordinate concepts, subject, predicate, and object, causality, contrast, and inclusion), and can be generated from the statistical characteristics of the surrounding context. For example, words such as "electrically assisted bicycle" and "urban transportation" are tagged with relational tags that designate the former as a subordinate concept of transportation and the latter as a superordinate concept of transportation. Relational tags are associated with each word in the decomposed semantic units and registered in the semantic unit dictionary (main body), which will be described later.

[0035] Each time a new document is input, the determination unit 140 updates the average similarity or variance value of the similar concept cluster and determines whether to absorb the word resolved from the document into an existing similar concept cluster or to generate a new similar concept cluster related to the word. A similar concept cluster is a grouping of multiple words with similar meanings or characteristics into one cluster. The similar concept cluster is registered as part of the semantic unit dictionary, which will be described later.

[0036] When a user document or a public document is input, the automatic generation unit 150 automatically generates training data for the document using a semantic unit dictionary. The training data is stored in the storage unit 300, and the semantic unit dictionary is learned and reconstructed based on the training data at a predetermined timing.

[0037] The output unit 200 outputs text generated by analyzing or translating the user document using the processing unit 220, which will be described later. For example, when a request for analysis or translation of the user document is made from the user terminal 3, the output unit 200 transmits the analyzed or translated text as response data to the user terminal 3 via the network.

[0038] The extraction unit 210 extracts character strings from user documents in the form of phrases, propositions, or logic units based on a semantic unit dictionary. A phrase is a sentence that is not a complete sentence but has a cohesive meaning. A proposition is a statement that can be determined to be true or false. A logic is a statement that links propositions together and refers to causal relationships or inferences.

[0039] Specifically, the extraction unit 210 references the similarity scores, the base propositions, and the relation tags registered in the semantic unit dictionary (described later) and classifies the strings from the user document as phrases, propositions, or logic, before extracting them. For example, for phrases, the extraction unit 210 extracts strings corresponding to the phrases by determining the similarity scores registered in the semantic unit dictionary. For propositions, the extraction unit 210 compares the logical structure of the subject, predicate, and object with the base propositions registered in the semantic unit dictionary and evaluates whether they match or mismatch, thereby extracting strings corresponding to the propositions. The base propositions are a set of mathematically and logically correct propositions that serve as the basis and foundation for analyzing and inferring a given sentence, and are registered as part of the semantic unit dictionary. For logic, the extraction unit 210 extracts strings corresponding to logic by determining the relationships between the strings related to the propositions based on the relation tags registered in the semantic unit dictionary.

[0040] More specifically, regarding phrases, the extraction unit 210 calculates the cosine similarity or embedding distance in a semantic space (vector space) as the similarity between words contained in the user document through arithmetic processing, and extracts character strings in phrase units based on the calculation results. Cosine similarity and embedding distance are parameter values ​​that measure the similarity in meaning between words, and the closer the meaning, the more likely the character strings are to be extracted as a phrase. For example, if the word "go to school" registered in the dictionary is "going to school," and the character string "going to school" in the input document is "going to school," the result is a partial match and a high similarity score, and the registered word "go to school" and the character string "going to school" are treated as the same phrase unit.

[0041] More specifically, regarding propositions, if a proposition such as "1+1=2" is registered in the set of reference propositions and a character string such as "1+1=3" is found in the user document, the extraction unit 210 will determine that the character string is structurally a proposition but is a contradictory proposition that is inconsistent with the set of reference propositions.

[0042] More specifically, when a user document contains a character string such as "When it rains, the ground gets wet," the extraction unit 210 extracts the character string as a logic tagged with a relation indicating "cause and effect."

[0043] The extraction unit 210 extracts character strings in units of phrases, propositions, or logic from user documents based on the similarity scores, reference propositions, and relation tags registered in the semantic unit dictionary, and provides the extracted character strings, to the processing unit 220 described below, as basic data to which labels indicating possible analysis or translation errors (error candidate labels) have been added. The error candidate labels are labels that indicate parts of the character string that cannot be determined to be errors with certainty but may be errors.

[0044] Furthermore, the extraction unit 210 refers to the similar concept clusters registered in the semantic unit dictionary, and extracts character strings that correspond to the most representative, highest ranking, or characteristic phrases included in these similar concept clusters.

[0045] The processing unit 220 performs processing for analyzing or translating the user document based on the character string extracted by the extraction unit 210. Specifically, when the processing unit 220 receives basic data including the extracted character string and the error candidate labels from the extraction unit 210, the processing unit 220 adds the error candidate labels from the basic data or corrects the user document based on the error candidate labels, and then generates text by analyzing or translating the user document. In other words, the processing unit 220 corrects the analysis accuracy or translation accuracy using the basic data. The generated text is displayed via the output unit 200 or returned to the user terminal 3. For example, when there is a character string that is structurally a proposition but is inconsistent with the reference propositions, the processing unit 220 performs correction processing to generate text in which the character string is labeled with an error candidate label.

[0046] The re-evaluation unit 230 re-evaluates the analyzed or translated text analyzed or translated by the processing unit 220 based on the semantic unit dictionary, points out or corrects any unanalyzable or mistranslated parts, and then regenerates the text. The re-generated text is also displayed via the output unit 200 or sent back to the user terminal 3.

[0047] The storage unit 300 stores an updatable semantic unit dictionary and also stores training data. Although not shown, the storage unit 300 also stores programs, data, etc. to be executed by the functional components of the processor 10. When executing the dictionary creation / update process or the document analysis / translation process described below, the semantic unit dictionary is read out along with these programs, data, etc.

[0048] [Semantic Unit Dictionary] A semantic unit dictionary is a database that maps the decomposed words generated by syntactically analyzing an input document to a semantic structure. Each word is associated with a similarity score, relational tag, location information, and context information. As shown schematically in Figures 3 and 4, a semantic unit dictionary is composed of a body of text, a set of base propositions, and a cluster of similar concepts. The similarity score is numerical data that evaluates the semantic agreement or similarity between words. The relational tag is a label that indicates the relationship between words or propositions, such as superordinate / subordinate concept, subject / predicate / object, causal, contrast, and inclusion. These relational tags can be used to determine whether an extracted string is classified as a phrase, proposition, or logic. The location information indicates the location of each word in the document. The context information indicates the context surrounding the location. The set of base propositions is a collection of propositions that have been confirmed to be mathematically and logically correct. By comparing them with the propositions in the input document, they serve as a basis for evaluating consistency and inconsistency. Similar concept clusters are data that group together concepts, words, etc. that are similar in meaning.

[0049] As shown in Figure 3, the semantic unit dictionary itself is a database that associates an input document with each word decomposed into the input document, each word mapped to a semantic structure, similarity scores between words, relational tags between words, and meta-information (location information and context information) for each word. The semantic unit dictionary itself is updated each time an input document is parsed.

[0050] As shown in Figure 4, the standard propositions are organized as a database divided into groups of propositions related to various phenomena. The standard propositions are constructed by obtaining them from a large-scale linguistic knowledge base such as WordNet.

[0051] As shown in Fig. 4 as an example, a similar concept cluster is configured as a database that brings together multiple words with similar meanings and characteristics. Each time an input document is parsed, if it is determined that a word decomposed from the document can be absorbed, a new word is added to the similar concept cluster, while if it is determined that a word decomposed from the document does not belong to any of the existing similar concept clusters, a new similar concept cluster for the word is generated.

[0052] [Dictionary creation and update process] Next, the dictionary creation and update process executed by the language processing computer 1 will be described with reference to FIG.

[0053] As shown in Fig. 5, in the dictionary creation and update process, first, a public document or a user document is input to the input unit 100 (S1). A public document is input, for example, in response to a request to start the dictionary creation and update process or at a predetermined regular interval, and a user document is input in response to a request from the user terminal 3. For example, the input document (input document) may be, "The spread of electrically assisted bicycles is changing transportation in urban areas."

[0054] Next, the parser 110 performs a parsing of the input document and breaks it down into elements of word units defined in the large-scale language knowledge base (S2). As a result, the broken down words become, for example, "electrically assisted bicycles," "popular," "urban mobility," and "changing."

[0055] Next, the mapping unit 120 maps each decomposed word to a semantic structure (semantic space) (S3). That is, each decomposed word is associated with a word in the semantic space and registered in the main body of the semantic unit dictionary. For example, each word mapped to the semantic structure is "electrically assisted bicycle," "expansion," "change in urban structure," and "change." At this time, the mapping unit 120 identifies position information of the appearance position of each word in the document and context information indicating the context surrounding the appearance position, and this position information and context information are also registered in the main body of the semantic unit dictionary as meta information for each word (see FIG. 3).

[0056] Next, the evaluation unit 130 generates a similarity score by evaluating the semantic agreement or similarity between the words based on the vector information in the semantic structure of each word, and generates a relation tag by evaluating the relationship between the words (S4). The generated similarity score and relation tag are registered in the main body of the semantic unit dictionary (see FIG. 3).

[0057] In this way, the main body of the semantic unit dictionary is created and updated, and the main body of the semantic unit dictionary is stored in the storage unit 300 (S5). Note that the storage unit 300 stores the base proposition groups and similar concept clusters in advance.

[0058] Next, the determination unit 140 determines whether or not to newly generate a similar concept cluster as a similar concept cluster related to the word registered in the semantic unit dictionary (S6). In this determination process, the average similarity or variance value of the similar concept cluster including the target word is updated, and if the difference between before and after the update is greater than a predetermined value, it is determined that a new similar concept cluster to which the target word belongs should be generated. On the other hand, if the difference is smaller than the predetermined value, it is determined that the target word should be absorbed (added) to an existing similar concept cluster. Note that if an existing similar concept cluster contains an identical word, it is desirable not to absorb it, since there will be no difference between before and after the update.

[0059] In S6, if it is determined that a new similar concept cluster should be newly generated (S6: YES), the determination unit 140 generates a new similar concept cluster to which the target word belongs, and stores the new similar concept cluster in the storage unit 300 (S7).

[0060] On the other hand, in S6, if it is determined that the target word should be absorbed (added) to the existing similar concept cluster without generating a new similar concept cluster (S6: NO), the determination unit 140 absorbs (adds) the target word to the existing similar concept cluster, updates the existing similar concept cluster, and stores it in the storage unit 300 (S8).

[0061] As described above, each time a document is input, the automatic generation unit 150 automatically generates training data for the document using the semantic unit dictionary (S9). The training data is input at a predetermined timing, an output result is predicted based on the semantic unit dictionary, an error is found by comparing this prediction with the training data, and the contents of the semantic unit dictionary are corrected based on this error using, for example, error backpropagation, thereby repeatedly learning and reconstructing the semantic unit dictionary using the training data.

[0062] [Document analysis and translation processing] Next, the document analysis and translation process executed by the language processing computer 1 using the above-mentioned semantic unit dictionary will be described with reference to FIG.

[0063] As shown in Fig. 6, in the document analysis and translation process, first, a user document is input to the input unit 100 in response to a request from the user terminal 3 (S20). For example, the input document may be "AI technology is improving diagnostic accuracy in the medical field."

[0064] Next, the extraction unit 210 extracts character strings that form phrase units, which are groups of words included in the input document, based on the similarity scores registered in the semantic unit dictionary (S21). At this time, the extraction unit 210 references the similar concept clusters in the semantic unit dictionary and also extracts character strings that correspond to the most representative, highest ranking, or characteristic phrases included in the similar concept clusters. As a result, phrase-based character strings such as "AI technology" → "artificial intelligence technology," "medical field" → "clinical field," "diagnostic accuracy" → "diagnostic accuracy," and "enhancing" → "improving" are extracted from the input document.

[0065] Next, the extraction unit 210 extracts character strings that form propositional units, which are groups of words contained in the input document, based on the base propositions registered in the semantic unit dictionary (S22). As a result, character strings that form propositional units, such as "AI technology is making clinical diagnoses more accurate," are extracted from the input document.

[0066] Next, the extraction unit 210 extracts character strings that form logical units of words contained in the input document based on the relational tags and meta-information (word position information and context information) registered in the semantic unit dictionary (S23). As a result, character strings of logical units such as "AI technology - diagnostic accuracy (clinical)" are extracted from the input document. Each character string extracted in this way by phrase, proposition, and logical unit is provided to the processing unit 220 as basic data with an error candidate label attached.

[0067] Next, the processing unit 220 performs analysis or translation processing on the character strings extracted in phrase units, proposition units, and logical units (S24). At this time, if there are character strings labeled with error candidate labels, the processing unit 220 also performs separate analysis or translation processing according to the error candidates.

[0068] Next, the output unit 200 outputs text that is the result of analyzing or translating the input document, and transmits it to the user terminal 3 or outputs (displays or prints) it on the output device 13 (S25). As a result, for example, text such as "Artificial intelligence technology is improving the accuracy of clinical diagnoses" is output.

[0069] Furthermore, the re-evaluation unit 230 re-evaluates the analyzed or translated output text based on the semantic unit dictionary, and points out or corrects unanalyzable or mistranslated parts (S26). As a result, if the text contains a character string labeled as an error candidate, the character string corresponding to the error candidate is extracted, analyzed, or translated again, and a different text with the errors pointed out or corrected is output.

[0070] For reference, FIG. 7 is a flowchart showing the flow of conventional token-based document analysis and translation processing. As shown in the figure, in token-based document analysis and translation processing, when a document is input (S30), words from the document are broken down into tokens (S31), and then analysis and translation processing is performed on each token (S32). Text is output by combining the analyzed and translated word strings for each token (S33). In such conventional token-based document analysis and translation processing, it is difficult to correct semantic errors in the output analyzed and translated text. On the other hand, in the document analysis and translation processing by the language processing computer 1 of this embodiment, text analyzed and translated on a semantic basis is output, so semantic error correction can also be appropriately performed by reevaluation using a semantic unit dictionary.

[0071] With this type of semantic unit analysis or translation, for example, an input document such as "AI technology is improving diagnostic accuracy in the medical field" is converted into text such as "Artificial intelligence technology is improving the accuracy of clinical diagnoses." In contrast, with conventional techniques that analyze and translate on a token-by-token basis, the same input document is broken down into smaller tokens, such as "AI technology is improving diagnostic accuracy in the medical field," and converted into text such as "Artificial intelligence technology is improving the accuracy of judgments in the medical field." Comparing the two texts from a natural language perspective, the former (semantic unit) is more accurate and easier to read than the latter (token unit) in terms of phrases, propositions, and logic. For example, with regard to phrases, the latter "artificial intelligence technology" is a relatively long phrase that undeniably feels mechanically generated, while the corresponding former "artificial intelligence technology" is more natural and readable due to the separation of the words "artificial intelligence" and "technology."

[0072] The language processing computer 1 described above can basically generate and infer natural language text by processing in semantic units, but in order to ensure compatibility with token-based analysis and translation processing, it may be connectable to existing token-model AI systems and may be equipped with an intermediate translation unit that converts input documents between them.

[0073] According to the language processing computer 1 of this embodiment, the accuracy of language analysis and translation can be improved by performing processing on a semantic basis, and the accuracy of language analysis and semantic analysis, which are the prerequisites for artificial intelligence learning, can be improved.

[0074] According to the dictionary creation / update process and document analysis / translation process, since they are replaced with processes on a semantic basis, it is possible to improve the conversion accuracy and semantic consistency between multiple languages.

[0075] To give a specific example, by using language processing computer 1 when processing translations, dialogues, contracts, etc., it is possible to obtain natural, consistent content that has been analyzed and translated at semantic units, thereby preventing misunderstandings. For example, in fields where mistranslations are unacceptable, such as law, finance, medicine, and education, if language processing computer 1 is used to analyze and translate contracts, medical records, etc. as input documents, it can obtain higher evaluation results than existing AI that processes at token levels, based on indicators such as logical consistency, nuance retention rate, and mistranslation reduction rate. In other words, it is possible to improve the explainability of AI (XAI: Explainable AI), which is the ability of AI to explain how it made decisions and predictions in a way that is understandable to humans.

[0076] It can also be applied to pre-learning processing by AI, and can be used professionally, for example, in the education industry or in the development stage of translation models.

[0077] Furthermore, the functional configuration of language processing computers 1 can be incorporated as part of AI functions, dramatically improving the "understanding" and "explainability" of AI. For example, this will raise the level of practical application of AI to areas that require high accuracy to prevent expression errors, such as legal, medical, educational, and administrative documents.

[0078] The present invention is not limited to the above-described embodiment, and can be modified as appropriate within the scope of the present invention.

[0079] Some or all of the above-described embodiments can also be described as follows:

[0080] (Appendix 1) a parsing unit that performs a syntactic analysis of an input document and divides the document into elements of each word; a mapping unit that maps each of the decomposed words to a semantic structure; an evaluation unit that evaluates the semantic agreement or similarity between the words divided into the word-unit elements; a storage unit for storing a semantic unit dictionary as a database in which the scores indicating the similarity between the words are mapped to a semantic structure by associating each of the words, the evaluated scores, and the documents, and relational tags indicating the relationship between the words or propositions, including causation, inclusion, contrast, subject-predicate-object relationship, and superordinate-subordinate relationship, are recorded; an extraction unit that extracts character strings from a predetermined document in units of phrases, propositions, or logic based on the semantic unit dictionary; a processing unit that processes analysis or translation of the predetermined document based on the extracted character string; A language processing computer comprising:

[0081] (Appendix 2) The language processing computer described in (Appendix 1), wherein the semantic unit dictionary includes the relationship tags that indicate the relationships between the words divided into the word-unit elements, and a set of reference propositions that are pre-registered to ensure logical accuracy and serve as standards for determining consistency or contradiction with the propositions in the specified document.

[0082] (Appendix 3) The language processing computer according to (Appendix 2), wherein the extraction unit refers to the score, the set of reference propositions, and the relation tag contained in the semantic unit dictionary, and extracts the character string after classifying it into one of the phrase, the proposition, or the logic.

[0083] (Appendix 4) The extraction unit extracting the character string related to the phrase by determining the phrase based on the score included in the semantic unit dictionary; extracting the character strings related to the proposition by comparing the logical structure of the subject, predicate, and object of the proposition with the reference propositions included in the semantic unit dictionary and evaluating whether they match or not; The language processing computer described in (Appendix 2) extracts the character strings related to the logic by determining the relationship between the multiple character strings related to the proposition based on the relation tags included in the semantic unit dictionary.

[0084] (Appendix 5) the extraction unit extracts character strings from the predetermined document in units of the phrase, the proposition, or the logic based on the scores, the standard propositions, or the relation tags contained in the semantic unit dictionary, and provides the extracted character strings as basic data to which labels indicating possible analysis or translation errors have been added; The language processing computer according to (Supplementary Note 2), wherein the processing unit corrects analysis accuracy or translation accuracy using the basic data.

[0085] (Appendix 6) A step of syntactically analyzing the input document and dividing the document into word-based elements; mapping each of the decomposed words to a semantic structure; a step of evaluating the semantic agreement or similarity between the words divided into the word-unit elements; a step of storing a semantic unit dictionary as a database in which the scores indicating the similarity between the words and the semantic structure by associating each of the words, the evaluated scores, and the documents are recorded together with relation tags indicating the relationship between the words or propositions, including causation, inclusion, contrast, subject-predicate-object relationship, and superordinate-subordinate relationship; Extracting character strings from a given document in units of phrases, propositions, or logic based on the semantic unit dictionary; analyzing or translating the predetermined document based on the extracted character string; A language processing method characterized by executing the above by a language processing computer.

[0086] (Appendix 7) A step of syntactically analyzing the input document and dividing the document into word-based elements; mapping each of the decomposed words to a semantic structure; a step of evaluating the semantic agreement or similarity between the words divided into the word-unit elements; a step of storing a semantic unit dictionary as a database in which the scores indicating the similarity between the words and the semantic structure by associating each of the words, the evaluated scores, and the documents are recorded together with relation tags indicating the relationship between the words or propositions, including causation, inclusion, contrast, subject-predicate-object relationship, and superordinate-subordinate relationship; extracting character strings in units of phrases, propositions, or logic from a given document based on the semantic unit dictionary; processing an analysis or translation of the predetermined document based on the extracted character string; A language processing computer readable program for executing the above.

[0087] (Appendix 8) The language processing computer according to claim 1, wherein the extraction unit calculates the similarity between words included in the specified document in a vector space and extracts a string of characters in units of a phrase based on the calculation result.

[0088] (Appendix 9) The language processing computer according to (Supplementary Note 1), wherein the mapping unit maps each of the words to an abstracted and normalized semantic structure based on a lexical knowledge base.

[0089] (Appendix 10) the mapping unit acquires, as meta-information, position information indicating a position of appearance in the document and context information indicating a context surrounding the position of appearance for each of the words mapped to the semantic structure; The language processing computer according to (Supplementary Note 1), wherein the semantic unit dictionary includes the meta-information corresponding to each of the words.

[0090] (Appendix 11) The semantic unit dictionary includes, as a similar concept cluster, a group of words that have a high semantic match or similarity between the words based on the evaluated score; the extraction unit refers to the similar concept cluster and extracts the character string based on the semantic unit dictionary; The language processing computer according to (Appendix 1), further comprising a determination unit that updates an average similarity or a variance value of the similar concept cluster each time a new document is input, and determines whether to absorb the word into the existing similar concept cluster or to generate a new similar concept cluster.

[0091] (Appendix 12) The language processing computer according to (Supplementary Note 1), further comprising an automatic generation unit that automatically generates training data for documents using the semantic unit dictionary.

[0092] (Appendix 13) The language processing computer described in (Appendix 1) further comprises a re-evaluation unit that re-evaluates the analyzed or translated text processed for the analysis or translation of the specified document based on the semantic unit dictionary, and points out or corrects unanalyzable or mistranslated parts.

[0093] (Appendix 14) The language processing computer according to (Appendix 1), wherein the processing unit is a processing system that performs natural language generation and inference based on a flow of semantic units, and generates and infers natural language by bypassing token processing.

[0094] (Appendix 15) A language processing computer as described in (Appendix 1), further comprising an intermediate compatibility unit that ensures connectivity with current token-based systems and converts between analysis or translation between semantic units and token units. [Explanation of symbols]

[0095] 1. Language processing computer 2 AI Server 3. User terminal 10 processors 11 Mass storage 12 Input Devices 13 Output Devices 14. Communications equipment 100 Input section 110 Parser 120 Mapping Section 130 Evaluation Department 140 Judgment section 150 Automatic generation section 200 Output section 210 Extraction part 220 Processing section 230 Reevaluation Department 300 Storage section

Claims

1. a parsing unit that performs a syntactic analysis of an input document and divides the document into elements of each word; a mapping unit that maps each of the decomposed words to a semantic structure; an evaluation unit that generates a similarity score by evaluating the semantic agreement or similarity between the words divided into the word-unit elements; a storage unit that stores a semantic unit dictionary as a database that records the scores that map each word to a semantic structure by associating each word with the generated scores and the documents, and the scores that indicate the similarity between the words, and relationship tags that indicate the relationship between the words or propositions and that include causation, inclusion, contrast, subject-predicate-object relationship, and superordinate-subordinate relationship; an extraction unit that extracts character strings from a predetermined document in units of phrases, propositions, or logic based on the semantic unit dictionary; a processing unit that processes analysis or translation of the predetermined document based on the extracted character string; A language processing computer comprising:

2. 2. The language processing computer of claim 1, wherein the semantic unit dictionary includes the relation tags indicating the relationships between the words divided into the word-unit elements, and a set of reference propositions that are pre-registered to ensure logical accuracy and serve as a basis for determining consistency or contradiction with the propositions in the specified document.

3. 3. The language processing computer according to claim 2, wherein the extraction unit refers to the score, the group of base propositions, and the relation tag included in the semantic unit dictionary, and extracts the character string after classifying it into one of the phrase, the proposition, or the logic.

4. The extraction unit extracting the character string related to the phrase by determining the phrase based on the score included in the semantic unit dictionary; extracting the character string related to the proposition by comparing the logical structure of the subject, predicate, and object of the proposition with the reference propositions included in the semantic unit dictionary and evaluating whether they match or not; 3. The language processing computer according to claim 2, wherein the character strings related to the logic are extracted by determining the relationship between the character strings related to the proposition based on the relation tags included in the semantic unit dictionary.

5. the extraction unit extracts character strings from the predetermined document in units of the phrase, the proposition, or the logic based on the scores, the standard propositions, or the relation tags included in the semantic unit dictionary, and provides the extracted character strings as basic data to which labels indicating possible analysis or translation errors have been added; The language processing computer according to claim 2 , wherein the processing unit corrects analysis accuracy or translation accuracy using the basic data.

6. A step of syntactically analyzing the input document and dividing the document into word-based elements; mapping each of the decomposed words to a semantic structure; generating a similarity score by evaluating the semantic agreement or similarity between the words divided into the word-unit elements; a step of storing a semantic unit dictionary as a database in which the scores indicating the similarity between the words and the semantic structure are mapped by associating each of the words with the generated scores and the documents, and relational tags indicating the relationship between the words or propositions, including causation, inclusion, contrast, subject-predicate-object relationship, and superordinate-subordinate relationship, are recorded; extracting character strings from a predetermined document in units of phrases, propositions, or logic based on the semantic unit dictionary; analyzing or translating the predetermined document based on the extracted character string; A language processing method characterized by executing the above by a language processing computer.

7. A step of syntactically analyzing the input document and dividing the document into word-based elements; mapping each of the decomposed words to a semantic structure; generating a similarity score by evaluating the semantic agreement or similarity between the words divided into the word-unit elements; a step of storing a semantic unit dictionary as a database in which the scores indicating the similarity between the words and the semantic structure are mapped by associating each of the words with the generated scores and the documents, and relational tags indicating the relationship between the words or propositions, including causation, inclusion, contrast, subject-predicate-object relationship, and superordinate-subordinate relationship, are recorded; extracting character strings from a predetermined document in units of phrases, propositions, or logic based on the semantic unit dictionary; processing an analysis or translation of the predetermined document based on the extracted character string; A language processing computer readable program for executing the above.

Citation Information

Patent Citations

  • Phrase-Based Dialogue Modeling Has Specific Use in Creating Recognition Grammar for Voice-Controlled User Interfaces

    JP2003505778A

  • Topic structure extracting method and device and topic structure extracting program and computer-readable storage medium with topic structure extracting program recorded thereon

    JP2005122510A

  • Dictionary creating device

    JP2007213336A

  • Analyzing method of character information, information analyzing device and program

    JP2013257756A

  • Indexing content at semantic level

    US20110196670A1